Skip to content

Pipeline Overview

Jill V. Hagey, PhD edited this page Feb 6, 2023 · 92 revisions

Pipeline Summary:

Pipeline Workflow


QC

  1. PhiX174 read removal and adapter removal using BBDuK

  2. Filtering, trimming, and base correction using fastp that includes:

    • quality trimming with a window size of 20 and quality of 30
    • quality pruning at 3' and 5' ends
    • removal of short reads
    • forced polyG tail trimming
  3. Contamination check of trimmed reads using Kraken2.

    • The kraken2 database that is used is the standard-8 one released by Ben Langmead, but YOU MUST download the 5/17/2021 version. If you used another kraken_db it is likely that you will have errors during the kraken step as files that are used in the PHoeNIx pipeline depend on matching the taxa IDs in the database.
    • For information on what species are included in the database see the inspect.txt file that accompanies the database on Ben's GitHub.

Analysis of Trimmed Reads

  1. QC Metrics Generated (all data generated for paired and unpaired reads generated post-trimming):
    • Number of total reads/bases
    • Percent of reads/bases remaining (from raw sequences)
    • Number of Q20/Q30 bases
    • Percent Q20/Q30 bases

Analysis using Trimmed Reads

  1. Gene detection and allele calling for antibiotic resistance (AR) srst2 in gene mode. We have curated an AR gene database that is a combination of three AR gene databases with redundancies removed and gene names standardized.
    • This step is only run with -entry CDC_PHOENIX
    • The curated AR gene database includes complete genes from these AR gene databases:
  2. Contamination is checked by using Kraken2 on the trimmed reads.
  3. srst2 MLST
    • This step is only run with -entry CDC_PHOENIX

Assembly

  1. Assembly of trimmed reads using SPAdes
  2. Filter reads to remove any scaffolds less than 500bp in length.

QC of Assembled Scaffolds >= 500bps

  1. Assess assembly quality using QUAST and custom scripts
  2. QC Metrics Generated:
  • Trimmed coverage (total trimmed bases / assembly length)
  • Assembly ratio (assembly size / median genome size of species)

Analysis of Assembled Scaffolds >= 500bps

  1. Assess genome assembly for completeness using BUSCO. This step is only run with -entry CDC_PHOENIX
  2. The mast distance is calculated from a pre-calculated sketch created with Mash and the top 20 best matches are passed into FastANI for increased speed in species ID.
  3. Calculate the average nucleotide identity (between genomes) using FastANI to determine species.
  4. Type multiple loci to characterize isolates of microbial species using MLST
  5. AR genes and hypervirulence genes are detected using GAMMA. We have curated an AR gene database that is a combination of three AR gene databases with redundancies removed and gene names standardized. Plasmid markers are detected with GAMMA-S.
  1. PROKKA is run on the scaffolds to generated a translated .faa file and an annotated .gff file, which will be passed to AMRFinder.
  2. AMRFinderPlus is run and the point mutations are reported in the Phoenix_Output_Report.tsv. The translated .faa and annotated .gff files from PROKKA are passed to AMRFinder as described in the AMRFinder documentation.
  3. In addition to running Kraken2 on the trimmed reads, KRAKEN2 is run on the weighted assembled scaffolds using the same database. This additional step allows us to check if any contamination made it into the assembly and this taxa call will be used if FastANI fails.
  • Kraken2 is also run in its normally on the scaffolds (non-weighted). This step is only run with -entry CDC_PHOENIX

Pipeline QC Checks

Notes on Evaluating Genome Assembly Quality

There are 3 "C"s we are concerned with when evaluating genome assemblies:

  • Contiguity: the size and number of contigs.
  • Completeness: the content of contigs, particularly the gene content.
  • Correctness: ordering and location of contigs.

Auto PASS/FAIL

Evaluating the quality of a genome assembly is more of an art than clear cut rules. The auto "PASS/FAIL" are metrics we deem to be the bare minimum quality standards and are:

  • >30x coverage
  • Assembly ratio stdev <2.58
  • Min assembly length >1,000,000bp

In addition to this information, staff should also consider other QC metrics (see more below), what species is being sequenced (some species complexes might have lower quality assemblies) and what you plan to do with the data. If there are particular metrics you are interested in then please submit a feature request for consideration.

WARNINGS

Warnings are defined as "out of line with what is expected and MAY cause problems downstream". The following will produce WARNINGS in the synopsis file:

  • <1,000,000 total reads for each raw and trimmed reads
  • % reads with Q30 average for R1 (<90%) and R2 (<70%)
  • >200 scaffolds
  • Checking that %GC content is within 2.58 stdev away from the mean %GC content for the species determined
  • Contamination check
    • >30% unclassified reads
    • Confirm there is only 1 genera with >25% of assigned reads

ALERTS

Alerts are defined as "something to note, but doesn't mean it's a poor-quality assembly". The following will produce ALERTS in the synopsis file:

  • No orphaned reads found after trimming
  • <10 reference genomes for species identified so no stdev for assembly ratio or %GC content calculated
  • >150x coverage

Clone this wiki locally