Releases: rec3141/papa2
Releases · rec3141/papa2
Release list
papa2 1.0.0
papa2 1.0.0 — the first stable release: byte-identical to R dada2 1.40, and substantially faster.
Parity
Verified byte-identical to R dada2 1.40 (R 4.6) on real MiSeq data (6- and 120-sample benchmarks) and targeted fuzz suites:
- Filtering — R's exact order of operations and boundary behaviour: truncLen counted from the original 5' end, truncQ after trimming, strict minQ gated on truncQ, paired orient.fwd read swapping, PhiX screening against the full circular genome with per-strand non-overlapping kmer counts, zero-output file removal. 252/252 filtered FASTQs byte-identical across the benchmarks.
matchIDsre-pairing (CASAVA field detection, chunk-boundary carry) andqualityTypeauto-detection (Phred+33/+64) included. - Denoising — identical ASVs, per-sample abundances and maps, including
pool=TRUE,pool="pseudo", andpriors. - Merging — per-pair
preferfrom denoised n0, rejects dropped unless requested. - Chimera removal — consensus, pooled and per-sample modes identical, with R's per-method defaults (1.5/2 consensus, 2/8 otherwise).
- Primer removal —
removePrimersport with IUPAC-aware matching and full-IUPAC reverse complement. - Taxonomy — genus assignments identical;
assign_taxonomy(seed=N)reproduces an R session'sset.seed(N)bootstrap stream via an exact port of R's RNG (residual differences equal R's own run-to-run noise from its unseeded tie-breaking — see #4). - Error matrices agree with R to ~1e-15 (LAPACK vs R's loess internals); all downstream results are byte-identical.
Known deliberate quirks replicated from upstream: #3, #4 (reported upstream as benjjneb/dada2#2210, benjjneb/dada2#2211).
Performance (32-core machine, real MiSeq data)
| workload | R dada2 (multithreaded) | papa2 | speedup |
|---|---|---|---|
| 6 samples end-to-end | 284 s | 156 s | 1.8× |
| 120 samples end-to-end | 1019 s | 512 s | 2.0× |
| filterAndTrim (single-thread) | 69 s | 10 s | 7× |
| assignTaxonomy, SILVA 138.1 × 2000 ASVs | 71 s | 15 s | 4.7× |
| assignTaxonomy, SILVA 138.1 × 20,697 ASVs | ~35 min | 50 s | ~40× |
Highlights: -O2 on the C core (~7×), vectorized numpy filtering with isal-accelerated gzip, OpenMP chimera and taxonomy paths, automatic worker×thread splitting (DADA2_CORES, DADA2_WORKERS, DADA2_OMP_THREADS), spawn-safe worker pools (loky), and parallel list APIs for derep_fastq and merge_pairs.
New
dada(pool=..., priors=...),combine_dereps,match_ref,set_num_threadsassign_taxonomy(seed=...), CLI--seedand working--threadsfastq_paired_filter(match_ids=..., quality_type=...)- New dependencies:
isal,loky
v0.1.0 — Initial release
Complete Python port of DADA2. See https://rec3141.github.io/papa2/