-
Notifications
You must be signed in to change notification settings - Fork 0
1. Usage
This page describes how to run PoODLE and prepare the required input files.
Before running PoODLE, we strongly recommend identifying clusters (genomic context groups) using CorGe+. PoODLE is designed for high-resolution analysis of pre-defined clusters, not for initial large-scale clustering. Running PoODLE on poorly defined or overly broad groups can:
- obscure true transmission signals
- reduce core genome size
- increase computational cost
- complicate epidemiological interpretation
CorGe+ is optimized for speed, scale, and screening, while PoODLE provides fine-grained, high-resolution analysis.
Together, they form a two-stage surveillance workflow:
| Step | Tool | Purpose |
|---|---|---|
| 1 | CorGe+ | Rapid screening and cluster detection |
| 2 | PoODLE | Detailed SNP, recombination filtering, and pangenome analysis |
Install the following software:
-
Nextflow(≥ 22.10.1) -
A container runtime:
-
Docker(recommended for local runs) -
Apptainer(recommended for HPC) Singularity
-
For Apptainer or Singularity, use a persistent shared cache directory so container images can be reused across runs and accessed by all compute nodes.
For Apptainer:
export NXF_APPTAINER_CACHEDIR=/path/to/shared/nextflow/apptainer/cacheFor Singularity:
export NXF_SINGULARITY_CACHEDIR=/path/to/shared/nextflow/singularity/cacheYou can also set these paths in your Nextflow configuration using apptainer.cacheDir or singularity.cacheDir.
For Apptainer and Singularity users, we recommend pre-pulling all required PoODLE container images before launching the workflow. This helps avoid issues caused by multiple Nextflow tasks pulling images concurrently, such as race conditions or incomplete cache files.
Helper scripts are provided:
For Apptainer:
export NXF_APPTAINER_CACHEDIR=/path/to/shared/nextflow/apptainer/cache
bash prefetch_poodle_containers_apptainer.shFor Singularity:
export NXF_SINGULARITY_CACHEDIR=/path/to/shared/nextflow/singularity/cache
bash prefetch_poodle_containers_singularity.shNote
NXF_APPTAINER_CACHEDIR and NXF_SINGULARITY_CACHEDIR control where Nextflow stores SIF images. Apptainer and Singularity also have their own OCI/layer caches, such as APPTAINER_CACHEDIR and SINGULARITY_CACHEDIR, which are mainly used while pulling or converting OCI images.
The pipeline requires a CSV samplesheet describing the samples to analyze.
Specify the file using:
--input samplesheet.csvEach row corresponds to one isolate/sample.
The samplesheet must contain 8 columns with the following headers.
| Column | Description |
|---|---|
sample |
Unique sample identifier. Spaces in sample names are automatically converted to underscores (_). |
fastq_1 |
Full path to FastQ file for Illumina QC trimmed short reads 1. File has to be gzipped and have the extension .fastq.gz or .fq.gz. |
fastq_2 |
Full path to FastQ file for Illumina QC trimmed short reads 2. File has to be gzipped and have the extension .fastq.gz or .fq.gz. |
annotation |
Full path to file with annotated genomes. File can be gzipped and have the extension .gff, .gff3,.gbk, .gb, .gbff. |
assembly |
Full path to assembled genome file. File can be gzipped and have the extension .fasta, .fa, .fna, .fas
|
cluster_id |
Custom cluster id. This entry will be identical for multiple samples from the same cluster. Spaces in cluter ids are automatically converted to underscores (_). |
species |
Custom bacterial species name. This entry will be identical for multiple samples from the same species. Spaces in sample names are automatically converted to underscores (_). |
reference |
Full path to assembled reference genome file. This must be identical for multiple samples from the same cluster. File can be gzipped and have the extension .fasta, .fa, .fna, .fas
|
Supported input types for variant calling
The pipeline supports three types of input data for variant calling:
| Input type | Required columns |
|---|---|
| Paired-end reads |
fastq_1, fastq_2
|
| Single-end reads | fastq_1 |
| Assemblies only | assembly |
If FASTQ files are not available, leave the columns blank.
Important
All columns must still be present in the CSV file.
Supported annotation types for pangenome profiling
Different annotation formats are supported by Panaroo for pangenome analysis.
If your annotations are not standard GFF3 files with embedded FASTA sequences, use the --annotation_format parameter to specify the correct format.
| Annotation file type | Description | --annotation_format |
|---|---|---|
| GFF3 | GFF3 file with embedded FASTA (e.g. Prokka/Bakta output) |
gff (default) |
| GFF3 + FASTA | GFF3 file without embedded FASTA, with a separate assembly FASTA file (e.g. NCBI RefSeq) | split_gff |
GenBank (.gb, .gbk, .gbff) |
GenBank flat file containing both annotation and sequence | genbank |
Important
All samples within a run must use the same annotation format.
Tip
If you downloaded annotations from NCBI, you likely need --annotation_format split_gff or genbank.
sample,fastq_1,fastq_2,annotation,assembly,cluster_id,species,reference
SAMPLE_1,/path/S1_R1.fastq.gz,/path/S1_R2.fastq.gz,/path/S1.gff,/path/S1.fasta,HC1-C1,Escherichia_coli,/path/ref1.fasta
SAMPLE_2,/path/S2_R1.fastq.gz,/path/S2_R2.fastq.gz,/path/S2.gff,/path/S2.fasta,HC1-C1,Escherichia_coli,/path/ref1.fasta
SAMPLE_3,/path/S3.fastq.gz,,/path/S3.gff,/path/S3.fasta,outbreak_A,Pseudomonas_aeruginosa,/path/ref2.fasta
SAMPLE_4,,,/path/S4.gff,/path/S4.fasta,outbreak_A,Pseudomonas_aeruginosa,/path/ref2.fastaAn example samplesheet is available in assets/samplesheet.csv
nextflow run MDHHS-Bioinformatics/poodle \
-profile singularity \
--input samplesheet.csv \
--outdir poodle_resultsThis will execute the core analysis workflow which include variant calling and pangenome analysis.
Note
This command downloads this pipeline to ~/.nextflow/assets/MDHHS-Bioinformatics/poodle. You can download the pipeline in a different location using git clone https://github.com/MDHHS-Bioinformatics/poodle.git. To run the pipeline, specify the path to the cloned repository (e.g. nextflow run /path/to/poodle ...).
Example enabling optional analyses (recombination filtering with Gubbins and MashTree) and adjusting resources:
nextflow run MDHHS-Bioinformatics/poodle \
-profile singularity \
--input samplesheet.csv \
--outdir poodle_results \
--gubbins \
--mashtree \
--max_memory 50.GB \
--max_cpus 16 \
--max_time 8.hGubbins may fail when organisms are highly similar and no recombination is identified. Therefore, Gubbins errors are ignored. In these cases, you may see the following message:
-[MDHHS-Bioinformatics/poodle] Pipeline completed successfully, but with errored process(es)-
[xxxx/yyyyy] NOTE: Process `POODLE:GUBBINS (<species>_<cluster>)` terminated with an error exit status (1) -- Error is ignoredIf this occurs, we recommend rerunning PoODLE without --gubbins and using -resume to generate the report.
The pipeline produces the following directories:
work/ # Nextflow working directory
<outdir>/ # Final pipeline outputs
.nextflow.log # Execution log
The work/ directory contains intermediate files and may be deleted after successful completion.
-
Use high-quality sequences: Ideally, assemblies should have <500 contigs ≥500 bp, reads ≥30× Illumina coverage, and no contamination. Pipelines like
PHoeNIX,BactopiaandTheiaProkprovide quality checks. -
Disk cleanup: After the pipeline completes, you may safely remove the Nextflow
work/directory to reclaim space. -
Prefer reads over assemblies for SNP analysis
-
Using a reference from within the cluster is strongly recommended to avoid core genome shrinkage
-
Run with
--gubbinsfor highly recombinant species -
Interpret SNP thresholds in epidemiological context, not in isolation
-
A genomic cluster should contain > 4 closely related samples. We strongly recommend using PoODLE after
CorGe+, since CorGe+ identifies genomic context groups at different thresholds.
For reproducible analyses, run a specific pipeline release:
nextflow run MDHHS-Bioinformatics/poodle \
-r v1.0.0 \
-profile singularity \
--input samplesheet.csv \
--outdir resultsUsing version tags ensures the same pipeline code and container versions are used.
Nextflow caches pipeline code locally.
To update to the latest version:
nextflow pull MDHHS-Bioinformatics/poodle