Repository navigation
[3.3.0] - 2026-10-08
The shallowest library no longer has to decide which sites the whole cohort keeps. Until now one pool below vcffilter.minDP removed a site for every pool. vcffilter.keepLowDepthAsZero keeps the site instead and writes that pool's reads there as unread. It is off by default, so nothing changes until you turn it on.
Three things to read before upgrading, all under Changed: the default of variantCall.scaleMapQ moved from 50 to 100, and a migrated parameters.config keeps its 50; init_multi and uninstall_all are now two words; and every table, plot and report an analysis publishes is renamed, so anything that reads one by name needs the new name.
As with every release, results belong to the release that produced them: 3.3.0 neither continues nor analyzes a project that 3.2.0 ran. Keep 3.2.0 installed beside it for such a project, or reset the project and run it again under 3.3.0.
Added
vcffilter.keepLowDepthAsZeroandvcffilter.minSamples. With the switch on, every cell belowminDP(one pool's reads at one site) is written as unread, which the depth table carries as zeros and the frequency table asNA, and a site is kept when at leastminSamplesof its cells reachminDP.minSamplesdefaults to 2 and runs from 1 to the number of pools with reads; anything else is refused before the run starts. Every read still counts toward deciding that an allele exists, because the false-positive filter runs first; only the cells at the floor count toward measuring its frequency. The manual's depth and quality filter section has the details, including what it does toTOTAL_AD.- Settings that would make a run produce nothing are refused at step 0, and
PoolSeqFlow check projectreports the same findings before you run. AscaleMapQbelowvarQualMindiscards every read, asampleThresholdabove 1 removes every site, aploidyorpoolSizebelow 1 breaks every detection limit, and afastqc.memorywith a unit is refused by FastQC. Settings that work but cost more than they look, or change what a number means, are warned about and the run continues. The manual lists them all. - Files saved on Windows. A
metadata.csvorruns.csvthat Excel saved as "CSV UTF-8" is read as it is, and a UTF-16 one is refused with what to save it as instead. Aparameters.configthat begins with a byte-order mark is named bycheck project, with the command that removes it; Nextflow on its own refuses the file over a character nothing displays. - Native zsh completion, with a description beside each command. The installer prints the one line
~/.zshrcneeds. PoolSeqFlow analysis modules install allinstalls every published module this release can run that you do not have yet. Every release starts with an empty module store, so this is how yours come back after an upgrade. A module that will not install is named and does not stop the others.
Changed
variantCall.scaleMapQdefaults to 100, where it was 50. It is the-Cthat lowers the mapping quality of reads with many mismatches. Measured on real pools, 100 removes the sites that are artifacts (Ti/Tv 0.848, nearly four times the soft-clip bias of the sites kept), while 50 removed five times as many and took real variants with them. 50 also lowered the frequencies it kept, by 0.035 on average against 0.008 at 100, and by 0.081 in pools between 0.50 and 0.75. A migratedparameters.configkeeps 50:migrate_configcarries your value across, as it does every setting that still exists, and reports it underKept your value. Set it to 100 before the first run under 3.3.0. The manual has the full measurements.init_multiisinit multi, anduninstall_allisuninstall all. The old spellings stop with a message giving the new one.- Every table, plot and report an analysis publishes begins with
outputPrefix, a new setting at the top ofparameters.config. WithoutputPrefix = 'Test', basicstats publishesTest_depth.tsvwhere it publisheddepth.tsv. The scripts that produced a result keep their own names, as doREADME.md,CITATIONS.mdandreferences.bib, and the report isTest_basicstats_report_20261008-142233.pdf, stamped with the local time it was published, where it wasreport.pdf.vcf.fileNamefollowsoutputPrefixunless you name the VCF yourself, andmigrate_configsetsoutputPrefixto the VCF name your project already has, so nothing the pipeline publishes changes its name. Either name must begin with a letter or a digit and hold only letters, digits,.,_and-; anything else is refused before anything is computed, inparameters.configor in a run table. Choose the prefix before the first run: likevcf.fileName, it is recorded with the results, and step 0 refuses a change to it afterwards. - A pool with no reads at a site is
NAin the frequency table, where it was 0, because a frequency over zero reads does not exist. By default no such cell reaches a table: the depth filter now also drops a site with an unread pool atminDP0, which 3.2.0 kept and published as 0. WithkeepLowDepthAsZeroon,NAis what an unread cell reads. - Every step writes its log as it runs. A step that fails, or is killed, leaves what it had reached in
Logs/. Before this, a step copied its log there only at the end, so a failure lost exactly the log that would have explained it. Each attempt starts with a line naming the run, the session and the time. - Capping is about seven times faster on real data, with identical output.
- A call set or filter that leaves nothing stops the run and says which settings to look at, instead of finishing as though it had succeeded. Nothing is published from it.
- The website follows releases. It was rebuilt whenever the manual changed on
main, so it could describe a version that had not been released. It now deploys when a release is published, and when a module is published between releases. - Environments. FastQC 0.12.1 to 0.13.0. In the analysis environment, Perl 5.32.1 to 5.44.0, pandoc 3.11 to 3.12.1, R's future and doFuture to their next minor versions, and ggrepel 0.9.8 added for mds. Patch updates beneath them in both. bwa, samtools, bcftools, cutadapt, Trim Galore and snpEff did not move.
Fixed
- Every analysis report from 3.0.0 to 3.2.0 was empty.
report.pdfheld a header and a list of the folder's files: the report never found the folder it described, so it skipped every section. Each module now lays out its own, and basicstats, association and mds show their tables and plots. The report and the folder'sREADME.mdnamed the release as "unknown" andCITATIONS.mdnamed none; all three name the one that ran. - Installing a release again over itself took packages out of the analysis environment whenever a module in the store declared them too, and every module declares several of the release's own: ggplot2, future and Rcpp among them. Only what a module added beyond the release's environment comes out now.
- A module whose install was refused left the libraries it brought in the store. They now go with it.
uninstall allsaid everything was removed even when an installation could not be, and said there was nothing to do once no conda environment was left, whatever installations and wrappers remained. It removes those too, and names whatever it could not remove.
Analysis modules
These versions need 3.3.0's analysis environment. On an earlier release, analysis modules install passes them over and installs the newest version that release can run.
- association
20261008.003. Two ways it publishedperm_p0, which no permutation p can be. A site where any unit had no reads came out 0, andfdr_pwith it; in 3.2.0 that tookminDP0. And the strongest sites in a table, far beyond any rearrangement, failed to count themselves and came out 0 too. Each site is now tested over the units it was read in, and not tested when fewer than three were read.design_flooris 1/n!, where it said 2/n!, so four units can reach 0.042. Each phenotype is published in two files of its own,association_<phenotype>.tsvandalleles_<phenotype>.tsv, whereassociation.tsvandassociation_alleles.tsvheld every phenotype beside aphenotypecolumn. Anything but a letter, a digit,.,_or-in a phenotype's name becomes_, and two phenotypes that would share a file are refused before anything is fitted.chromosomesrestricts the fit to the sequences it names, as it did in 3.2.0, where the manual said it only chose which get a Manhattan plot: a run that names some adjustsfdr_pover their sites alone. A depth table of a single site no longer stops it. - mds
20261008.003. A pair of pools averaged over fewer than 30 shared sites is flagged, indistance.tsv, on the console and under the plot. In simulation, a distance near 0.02 measured over 30 shared sites is off by about half itself, and real sites are noisier. The pools' names onmds.pngare placed clear of one another, where they could print on top of each other, andanalysis.modules.mds.labels = falseleaves them off. - basicstats
20261008.003. A pool's depth summaries cover only the sites it was read at, anddepth.tsvanddiversity.tsvcount the rest in a newunmeasuredcolumn. A site with no reads for a pool was averaged in as a depth of 0, and one such site was enough to turn that pool's harmonic depth into 0 and its effective sample size into NA; in 3.2.0 that tookminDP0.unmeasuredis the fourth column of each, aftersites, so a script reading them by position needs updating. A run whose indel table is empty no longer stops it.
Commits
- (2a15aea) small change in releasing, test for linting
- (b96bfca) zsh completion added
- (c5bacb4) dos2unix for parameters, metadata, and runs files
- (51b649c) manual and documents passes
- (8d0447c) dropZeroDept toggle added
- (b6e0ce6) Comments shortened
- (ebd69af) Comments shortened
- (eedf598) scaleMapQ decisions are recorded
- (772db1f) Analysis modules are updated to be compatible with the zero depth sites change
- (85eb380) cap_depth is fixed for performance
- (bffc6a0) Empty vcf, depth, and freq files now produce error
- (ad85fb1) Logging reworked
- (aaeab55) Test suite for dropZeroDepth added
- (6746280) major allele to ref tests
- (f54a5ae) SampleID RG_Sample clarifications
- (8e8a078) parameter checks added to step0
- (9b80793) config migrate is updated after dropZeroDepth
- (6127b53) Full release cycle rework
- (b96b818) wrapper tidiness
- (2b50c1c) Releasing improvements
- (87edc97) releasing cycle streamlined further
- (f2c8187) Manual rewrite
- (998c307) prep version is rescoped
- (1224341) Minor fix on how prep version works
- (d76ba2f) Prep for v3.3.0
- (7419281) module publishing in release cycle is handled with a script now
- (1e36baf) mds site diversity limit established
- (e87e742) factorial calculation fixed
- (e04215b) dropZeroDepth replaced with keepLowDepthAsZero, and minSamples is added to vcfFilter
- (30aa234) Prep for v3.3.0
- (babb475) Version bump v3.3.0
- (6550cb5) releasing issues found, publish module updated
- (d8b7cee) Version bump reverted
- (1b04cb1) module updates, association analysis improved, report.pdf now shows the reports
- (86bb11b) Filename fixes for modules
- (60d6000) Large commit on releasing, modules install all, and test suite prep for 3.3.0
- (299d6e6) release fixes
- (c22acdb) Release prep for 3.3.0: refreshed environments, six module versions bumped, bump-analysis-version.sh --pending, clean-release-scratch.sh
Download and install
curl -LO https://github.com/ozankiratli/PoolSeqFlow/releases/download/v3.3.0/PoolSeqFlow-3.3.0.tar.gz
tar -xzf PoolSeqFlow-3.3.0.tar.gz
cd PoolSeqFlow-3.3.0
cp parameters.config.template parameters.config
./PoolSeqFlow installThen edit parameters.config and metadata.csv for your data and run
./PoolSeqFlow run.
Verify the download with sha256sum -c SHA256SUMS.
PoolSeqFlow.tar.gz is the same archive under a stable name, for
scripted installs:
https://github.com/ozankiratli/PoolSeqFlow/releases/latest/download/PoolSeqFlow.tar.gz
Upgrading an existing project? Your parameters.config is not
touched by a new version and can be missing parameters this release
expects. Run ./PoolSeqFlow migrate_config and read what it reports --
see Upgrading.
Full documentation: https://ozankiratli.github.io/PoolSeqFlow/
The full changelog, including every commit, is in CHANGELOG.md in the download and in the repository.