Skip to content

New pipeline: nf-core/plasmodiumpopgen #161

Description

@kathrynmurie

Pipeline title/name

plasmodiumpopgen

Keywords

genomics, microhaplotype, population, IBD, relatedness, plasmodium

What is it about?

This pipeline will provide a standardized workflow for population-genetic analysis of Plasmodium microhaplotypes generated from targeted sequencing data.

The pipeline will estimate commonly used metrics for Plasmodium, pop gen including:

  • Complexity of infection (COI)
  • Proportion of polyclonal infections
  • Between-host relatedness (IBD)
  • Within-host relatedness
  • Allele frequencies
  • Per-target summary

Per-target summary statistics will include:

  • Total and unique allele counts
  • Singleton alleles
  • Nucleotide diversity
  • Number of segregating sites
  • Tajima's D

The pipeline will support interoperable tools that users can select through command-line parameters, allowing methods to evolve as the field develops.

Plasmodium targeted sequencing presents a distinct pop gen problem because blood-stage parasites are haploid, while infections are frequently polyclonal and therefore contain unphased mixtures of multiple parasite genomes. Core quantities such as COI, within-host relatedness, allele frequencies, and between-host IBD require methods that explicitly account for this structure.

Although other eukaryotic pathogens can share individual biological features such as recombination or mixed infections, the combination of methods, defaults, benchmarking, and interpretation in this pipeline is designed and validated specifically for Plasmodium. Several core tools, including MOIRE and Dcifer, were developed specifically for inference from polyclonal malaria infections.

The pipeline design and method selection have been guided by the Plasmodium genomics community through the PlasmoGenEpi Network, including extensive benchmarking of methods.

Please provide a schematic diagram of the proposed pipeline

Image

What would a minimal first release of this pipeline include?

The initial release would:

  1. Split samples into user-defined populations for parallel analysis.
  2. Estimate COI using a naive estimator or MOIRE.
  3. Optionally...
    a. Estimate within-host relatedness and allele frequencies using MOIRE.
    b. Estimate between-host relatedness using Dcifer.
  4. Generate per-locus statistics using PGEcore.
  5. Collate results into standardized outputs.

This would provide an end-to-end workflow from processed targeted-sequencing/microhaplotype data to the principal metrics currently used for population-level analysis of Plasmodium.

I confirm my proposed pipeline will follow nf-core guidelines. Most importantly, my pipeline will:

  • be built with Nextflow.
  • pass nf-core lint tests and use standardized parameters.
  • be community-owned and developed within the nf-core organization.
  • open source under the MIT license with proper credits and acknowledgments.
  • have a descriptive, all lowercase, and without punctuation name.
  • use the nf-core pipeline template and predominantly use official nf-core modules.
  • focus on a specific data/analysis type with appropriate scope.
  • have properly maintained documentation.
  • be bundled using versioned Docker/Singularity containers.

Why do we need a new pipeline?

These analyses are commonly performed across the Plasmodium community, but there is currently no standardized nf-core workflow for running them consistently and at scale. Researchers instead have to connect individual tools using custom scripts, making analyses harder to reproduce and compare across studies.

The proposed pipeline does not introduce new methodology. Its value is in integrating established Plasmodium methods into a reproducible, scalable, community-maintained workflow with standardized inputs and outputs.

Importantly, this is not a generic population-genetics workflow applied to Plasmodium. A sample from a malaria infection can contain multiple genetically distinct clones. The observed alleles therefore represent a mixture of parasites, and important properties of those parasites must be inferred without necessarily being able to reconstruct the individual strains through phasing.

COI, allele-frequency estimation, within-host relatedness, and between-host IBD are consequently interconnected inference problems. The pipeline brings together methods developed specifically to make these population-level inferences from polyclonal Plasmodium infections.

There is currently no nf-core pipeline designed for this type of population-level analysis of Plasmodium targeted sequencing data.

Who would be interested?

Researchers working with Plasmodium targeted sequencing data. We hope this would eventually be useful in public health, although substantial work still needs to be done to translate these metrics to interpretation.

What has been done so far

We already have a prototype implementing some of this workflow in a repo called plasmodiumtransmissionintensity. It currently uses a large shared Docker image, we haven't filled in many parts of the template, and we want to change the name to plasmodiumpopgen. Therefore, we think starting from a fresh nf-core template and migrating the existing functionality will be the simplest path forward.

We actually proposed this pipeline to nf-core a year or so ago, and the proposal was approved as sufficiently distinct to warrant a new pipeline. However, we never settled on a pipeline name, so a dedicated nf-core channel was never created.

In the meantime, we focused on plasmodiumdrugres. We now have a dedicated team with capacity to bring plasmodiumpopgen through to an initial release over the next couple of months. Given the time since the original proposal and a conversation on slack, we're putting in a new proposal.

The PlasmoGenEpi Network has also conducted substantial benchmarking of the methods included in the pipeline and we will be releasing a preprint describing the benchmarking in the next couple of months.

URL to existing work (if applicable)

https://github.com/PlasmoGenEpi/plasmodiumtransmissionintensity/tree/dev

Are there any similar existing nf-core pipelines?

plasmodiumdrugres

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions