Skip to content

Releases: PathoGenOmics-Lab/fstic

v1.0.1: correctness fixes and documentation site

Choose a tag to compare

@Paururo Paururo released this 28 Jul 10:06
f7fb046

Correctness release. Two review passes went through the readers and the
estimators, the second one aimed specifically at the cases where Fstic produced
a plausible-looking number rather than the right one.

Full documentation now lives at pathogenomics-lab.github.io/fstic,
including a tutorial that
runs four samples from raw VCFs to a tree.

⚠️ Results change

If you have published numbers from 1.0.0, these are the reasons they may not
reproduce.

Loci where neither sample of a pair had a record were given empty frequency
vectors instead of being read as "both are reference". Chord added √2 for every
such locus, so two identical samples came out at 1.41, 2.83 and upwards
depending on how many private sites the rest of the cohort contributed, and GST
took 1.0 into its denominator, dropping from 1.0 to 0.5 purely because an
unrelated sample joined the run. A pairwise distance is now a property of the
pair. Affects chord, gst, reynolds, nei.

Reference bases could come from the wrong contig. When a contig name did not
match the FASTA, the lookup silently fell back to whichever contig sorted first.
Measured Bray-Curtis 0.4 against 0.2 on the same data. A single-contig reference
still resolves any name; with several contigs, an unmatched name is now reported.
Affects all estimators when --reference is used.

Allele frequencies summing above 1, which split multi-allelic records
produce, drove 1 - Σp² negative and took the estimators out of range. One
locus was enough to give FST 1.638 and Rogers 1.063, both bounded by 1. Such
sites are now rescaled and counted. Affects all.

Reynolds used Nei's GST in place of θ rather than the Reynolds, Weir &
Cockerham (1983) coancestry coefficient. The two agree only at completely fixed
loci, so values were not comparable with adegenet, Arlequin or GenAlEx output
carrying the same name. Affects reynolds.

Output depended on --workers. rayon's reduction tree varies with thread
count and floating point addition is not associative, so a 4-core and a 16-core
machine produced different files for the same command. On 320k loci, -w 3 gave
91356.2127736697 where every other count gave ...699. Affects all, last digits
only.

--formula fst is unchanged. It sums per-locus GST and is unbounded, which
contradicted the README; the README was corrected instead.

Silent failures are now loud

Inputs used to vanish without a word: a gzipped VCF failed on its first line and
produced an empty sample with exit code 0, a sample whose variants were all
filtered lost its matrix row, and two files with the same basename merged into
one. An unparseable AD slipped past --min-alt-reads 50, because a parse
failure became "field absent" and a missing field cannot be filtered on.

Unreadable files, files with no parseable data lines, duplicate sample names,
malformed FASTA records and tables mixing chrom conventions are now errors.
Everything dropped for a legitimate reason is counted and reported before the
run summary.

Table input matches its documentation

Case-insensitive column names and percentage frequencies were promised by the
README and not implemented, and both failed silently: a Sample,Position,...
header dropped every row, then the run died with an unrelated message. Both work
now. Duplicate rows keep the first value, as the VCF reader always did.

Also

  • --pass-only for the VCF FILTER column
  • FREQ computed from AD/DP when absent, with range validation on the single
    path both sources feed into
  • 83 tests, up from 57. Several existing ones could not fail: three asserted the
    output contains 0.000000, which the zero diagonal satisfies on any matrix.
  • Output precision at 10 decimal places, so closely related pairs no longer
    round to a tie
  • Non-finite distances written as NA rather than inf

Full changelog: CHANGELOG.md · Diff: v1.0.0...v1.0.1

First Release: Version 1.0.0 🎉

Choose a tag to compare

@Paururo Paururo released this 12 Aug 15:37
67a4c7c

We are pleased to announce the first public release of fstic, a high-performance, Rust-based command-line tool for computing pairwise genetic distances from variant data.
This version provides a complete, production-ready implementation with multi-core parallelism and extensive filtering options.

✨ Key Features

  1. Multiple Distance Metrics: FST, GST, Jost’s D, Reynolds, Nei, Cavalli-Sforza chord, Rogers, and Bray-Curtis.
  2. Flexible Input: Accepts single-sample VCFs or pre-processed allele-frequency tables (.csv, .tsv, .tab), with support for file lists.
  3. Smart Filtering: Configurable thresholds for depth, allele frequency, and allele counts.
  4. Configurable Output: Optional normalization by number of loci (--normalize) and clear, N×N distance matrices.
  5. Optimized for Speed: Fully parallelized using all available CPU cores (customizable with --workers).
  6. Transparent Reporting: Live progress, ETA, and detailed logs of applied filters.

Full Changelog: https://github.com/PathoGenOmics-Lab/fstic/commits/v1.0.0