Skip to content

Releases: rec3141/papa2

papa2 1.0.0

Choose a tag to compare

@rec3141 rec3141 released this 19 Aug 17:30

papa2 1.0.0 — the first stable release: byte-identical to R dada2 1.40, and substantially faster.

Parity

Verified byte-identical to R dada2 1.40 (R 4.6) on real MiSeq data (6- and 120-sample benchmarks) and targeted fuzz suites:

  • Filtering — R's exact order of operations and boundary behaviour: truncLen counted from the original 5' end, truncQ after trimming, strict minQ gated on truncQ, paired orient.fwd read swapping, PhiX screening against the full circular genome with per-strand non-overlapping kmer counts, zero-output file removal. 252/252 filtered FASTQs byte-identical across the benchmarks. matchIDs re-pairing (CASAVA field detection, chunk-boundary carry) and qualityType auto-detection (Phred+33/+64) included.
  • Denoising — identical ASVs, per-sample abundances and maps, including pool=TRUE, pool="pseudo", and priors.
  • Merging — per-pair prefer from denoised n0, rejects dropped unless requested.
  • Chimera removal — consensus, pooled and per-sample modes identical, with R's per-method defaults (1.5/2 consensus, 2/8 otherwise).
  • Primer removal — removePrimers port with IUPAC-aware matching and full-IUPAC reverse complement.
  • Taxonomy — genus assignments identical; assign_taxonomy(seed=N) reproduces an R session's set.seed(N) bootstrap stream via an exact port of R's RNG (residual differences equal R's own run-to-run noise from its unseeded tie-breaking — see #4).
  • Error matrices agree with R to ~1e-15 (LAPACK vs R's loess internals); all downstream results are byte-identical.

Known deliberate quirks replicated from upstream: #3, #4 (reported upstream as benjjneb/dada2#2210, benjjneb/dada2#2211).

Performance (32-core machine, real MiSeq data)

workload R dada2 (multithreaded) papa2 speedup
6 samples end-to-end 284 s 156 s 1.8×
120 samples end-to-end 1019 s 512 s 2.0×
filterAndTrim (single-thread) 69 s 10 s 7×
assignTaxonomy, SILVA 138.1 × 2000 ASVs 71 s 15 s 4.7×
assignTaxonomy, SILVA 138.1 × 20,697 ASVs ~35 min 50 s ~40×

Highlights: -O2 on the C core (~7×), vectorized numpy filtering with isal-accelerated gzip, OpenMP chimera and taxonomy paths, automatic worker×thread splitting (DADA2_CORES, DADA2_WORKERS, DADA2_OMP_THREADS), spawn-safe worker pools (loky), and parallel list APIs for derep_fastq and merge_pairs.

New

  • dada(pool=..., priors=...), combine_dereps, match_ref, set_num_threads
  • assign_taxonomy(seed=...), CLI --seed and working --threads
  • fastq_paired_filter(match_ids=..., quality_type=...)
  • New dependencies: isal, loky

v0.1.0 — Initial release

Choose a tag to compare

@rec3141 rec3141 released this 30 Mar 10:55

Complete Python port of DADA2. See https://rec3141.github.io/papa2/