Pipeline title/name
plasmodiumpopgen
Keywords
genomics, microhaplotype, population, IBD, relatedness, plasmodium
What is it about?
This pipeline will provide a standardized workflow for population-genetic analysis of Plasmodium microhaplotypes generated from targeted sequencing data.
The pipeline will estimate commonly used metrics for Plasmodium, pop gen including:
- Complexity of infection (COI)
- Proportion of polyclonal infections
- Between-host relatedness (IBD)
- Within-host relatedness
- Allele frequencies
- Per-target summary
Per-target summary statistics will include:
- Total and unique allele counts
- Singleton alleles
- Nucleotide diversity
- Number of segregating sites
- Tajima's D
The pipeline will support interoperable tools that users can select through command-line parameters, allowing methods to evolve as the field develops.
Plasmodium targeted sequencing presents a distinct pop gen problem because blood-stage parasites are haploid, while infections are frequently polyclonal and therefore contain unphased mixtures of multiple parasite genomes. Core quantities such as COI, within-host relatedness, allele frequencies, and between-host IBD require methods that explicitly account for this structure.
Although other eukaryotic pathogens can share individual biological features such as recombination or mixed infections, the combination of methods, defaults, benchmarking, and interpretation in this pipeline is designed and validated specifically for Plasmodium. Several core tools, including MOIRE and Dcifer, were developed specifically for inference from polyclonal malaria infections.
The pipeline design and method selection have been guided by the Plasmodium genomics community through the PlasmoGenEpi Network, including extensive benchmarking of methods.
Please provide a schematic diagram of the proposed pipeline
What would a minimal first release of this pipeline include?
The initial release would:
- Split samples into user-defined populations for parallel analysis.
- Estimate COI using a naive estimator or MOIRE.
- Optionally...
a. Estimate within-host relatedness and allele frequencies using MOIRE.
b. Estimate between-host relatedness using Dcifer.
- Generate per-locus statistics using PGEcore.
- Collate results into standardized outputs.
This would provide an end-to-end workflow from processed targeted-sequencing/microhaplotype data to the principal metrics currently used for population-level analysis of Plasmodium.
I confirm my proposed pipeline will follow nf-core guidelines. Most importantly, my pipeline will:
Why do we need a new pipeline?
These analyses are commonly performed across the Plasmodium community, but there is currently no standardized nf-core workflow for running them consistently and at scale. Researchers instead have to connect individual tools using custom scripts, making analyses harder to reproduce and compare across studies.
The proposed pipeline does not introduce new methodology. Its value is in integrating established Plasmodium methods into a reproducible, scalable, community-maintained workflow with standardized inputs and outputs.
Importantly, this is not a generic population-genetics workflow applied to Plasmodium. A sample from a malaria infection can contain multiple genetically distinct clones. The observed alleles therefore represent a mixture of parasites, and important properties of those parasites must be inferred without necessarily being able to reconstruct the individual strains through phasing.
COI, allele-frequency estimation, within-host relatedness, and between-host IBD are consequently interconnected inference problems. The pipeline brings together methods developed specifically to make these population-level inferences from polyclonal Plasmodium infections.
There is currently no nf-core pipeline designed for this type of population-level analysis of Plasmodium targeted sequencing data.
Who would be interested?
Researchers working with Plasmodium targeted sequencing data. We hope this would eventually be useful in public health, although substantial work still needs to be done to translate these metrics to interpretation.
What has been done so far
We already have a prototype implementing some of this workflow in a repo called plasmodiumtransmissionintensity. It currently uses a large shared Docker image, we haven't filled in many parts of the template, and we want to change the name to plasmodiumpopgen. Therefore, we think starting from a fresh nf-core template and migrating the existing functionality will be the simplest path forward.
We actually proposed this pipeline to nf-core a year or so ago, and the proposal was approved as sufficiently distinct to warrant a new pipeline. However, we never settled on a pipeline name, so a dedicated nf-core channel was never created.
In the meantime, we focused on plasmodiumdrugres. We now have a dedicated team with capacity to bring plasmodiumpopgen through to an initial release over the next couple of months. Given the time since the original proposal and a conversation on slack, we're putting in a new proposal.
The PlasmoGenEpi Network has also conducted substantial benchmarking of the methods included in the pipeline and we will be releasing a preprint describing the benchmarking in the next couple of months.
URL to existing work (if applicable)
https://github.com/PlasmoGenEpi/plasmodiumtransmissionintensity/tree/dev
Are there any similar existing nf-core pipelines?
plasmodiumdrugres
Pipeline title/name
plasmodiumpopgen
Keywords
genomics, microhaplotype, population, IBD, relatedness, plasmodium
What is it about?
This pipeline will provide a standardized workflow for population-genetic analysis of Plasmodium microhaplotypes generated from targeted sequencing data.
The pipeline will estimate commonly used metrics for Plasmodium, pop gen including:
Per-target summary statistics will include:
The pipeline will support interoperable tools that users can select through command-line parameters, allowing methods to evolve as the field develops.
Plasmodium targeted sequencing presents a distinct pop gen problem because blood-stage parasites are haploid, while infections are frequently polyclonal and therefore contain unphased mixtures of multiple parasite genomes. Core quantities such as COI, within-host relatedness, allele frequencies, and between-host IBD require methods that explicitly account for this structure.
Although other eukaryotic pathogens can share individual biological features such as recombination or mixed infections, the combination of methods, defaults, benchmarking, and interpretation in this pipeline is designed and validated specifically for Plasmodium. Several core tools, including MOIRE and Dcifer, were developed specifically for inference from polyclonal malaria infections.
The pipeline design and method selection have been guided by the Plasmodium genomics community through the PlasmoGenEpi Network, including extensive benchmarking of methods.
Please provide a schematic diagram of the proposed pipeline
What would a minimal first release of this pipeline include?
The initial release would:
a. Estimate within-host relatedness and allele frequencies using MOIRE.
b. Estimate between-host relatedness using Dcifer.
This would provide an end-to-end workflow from processed targeted-sequencing/microhaplotype data to the principal metrics currently used for population-level analysis of Plasmodium.
I confirm my proposed pipeline will follow nf-core guidelines. Most importantly, my pipeline will:
Why do we need a new pipeline?
These analyses are commonly performed across the Plasmodium community, but there is currently no standardized nf-core workflow for running them consistently and at scale. Researchers instead have to connect individual tools using custom scripts, making analyses harder to reproduce and compare across studies.
The proposed pipeline does not introduce new methodology. Its value is in integrating established Plasmodium methods into a reproducible, scalable, community-maintained workflow with standardized inputs and outputs.
Importantly, this is not a generic population-genetics workflow applied to Plasmodium. A sample from a malaria infection can contain multiple genetically distinct clones. The observed alleles therefore represent a mixture of parasites, and important properties of those parasites must be inferred without necessarily being able to reconstruct the individual strains through phasing.
COI, allele-frequency estimation, within-host relatedness, and between-host IBD are consequently interconnected inference problems. The pipeline brings together methods developed specifically to make these population-level inferences from polyclonal Plasmodium infections.
There is currently no nf-core pipeline designed for this type of population-level analysis of Plasmodium targeted sequencing data.
Who would be interested?
Researchers working with Plasmodium targeted sequencing data. We hope this would eventually be useful in public health, although substantial work still needs to be done to translate these metrics to interpretation.
What has been done so far
We already have a prototype implementing some of this workflow in a repo called plasmodiumtransmissionintensity. It currently uses a large shared Docker image, we haven't filled in many parts of the template, and we want to change the name to plasmodiumpopgen. Therefore, we think starting from a fresh nf-core template and migrating the existing functionality will be the simplest path forward.
We actually proposed this pipeline to nf-core a year or so ago, and the proposal was approved as sufficiently distinct to warrant a new pipeline. However, we never settled on a pipeline name, so a dedicated nf-core channel was never created.
In the meantime, we focused on plasmodiumdrugres. We now have a dedicated team with capacity to bring plasmodiumpopgen through to an initial release over the next couple of months. Given the time since the original proposal and a conversation on slack, we're putting in a new proposal.
The PlasmoGenEpi Network has also conducted substantial benchmarking of the methods included in the pipeline and we will be releasing a preprint describing the benchmarking in the next couple of months.
URL to existing work (if applicable)
https://github.com/PlasmoGenEpi/plasmodiumtransmissionintensity/tree/dev
Are there any similar existing nf-core pipelines?
plasmodiumdrugres