Releases: rnabioco/fqxv
Release list
v0.7.0
Minor: this cycle makes decode parallel-first — the archive is an analysis
substrate, not just cold storage. Additive Rust and Python APIs (stream
selection, a parallel record visitor, FASTA output) plus one
writer-side compatibility note: long-read archives written at the default
level now use chunked quality coding (context modes 5/6), which v0.6.x readers
refuse with an "upgrade fqxv" error. The on-disk format stays 1.0 — this
release reads every archive any earlier release wrote, and short-read archives
are byte-identical in both directions.
Added
- Stream-selective decode (#269):
StreamSelectionplus
decompress_records_select/RecordReader::with_selectionskip deselected
streams without entropy-decoding them. The range coders are ~96–98% of
decode compute, so sequence-only decode of an ONT archive runs ~12× faster.
decompress_fastaandfqxv decompress --fastaemit names+sequence as
single-line FASTA, skipping quality entirely. No format change; a full
selection keeps the canonical full-decode path bit-for-bit. - Parallel record visitor (#271, #276):
decompress_records_par— a
map-reduce visitor (init/visit/merge) that scales record consumers with
the block parallelism instead of funneling through one serial channel — and
decompress_records_par_select, its composition with stream selection
(parallel sequence-only decode measured 13× on ONT). - Python:
streams=andfasta=(#274):fqxv.open(..., streams="seq")
decodes only the selected streams;decompress_to_path/
decompress_to_bytesacceptfasta=True;fqxv.remotepasses both through. - Decode thread-scaling benchmark (#270):
bench/scripts/decode_scaling.sh
plusdocs/decode-scaling.mdmeasure full-decode throughput versus thread
count against gzip/pigz/zstd, with byte-identity asserted for every cell. Params::block_seq_bytes(#275): the long-read per-block raw-sequence
budget is a library parameter (0 = platform auto) instead of a hard-wired
constant.
Changed
- Long-read quality is coded in parallel-decodable chunks by default
(#279; design and measurements in #277/#278). Quality — 92–98% of long-read
decode — was one adaptive coding pass per block; it is now a serially-coded
warmup prefix plus K−1 whole-read chunks that encode and decode in parallel
(context modes5/6, K = 8, Nanopore + PacBio only). Measured on ONT/HiFi:
full decode ~1.9× faster at 16–32 threads (ONT 17.7 → 9.1 s, HiFi 66.5 →
33.8 s), HiFi compress 17% faster, for +0.36% / +0.19% archive size.
Compatibility: v0.6.x binaries cannot read mode-5/6 archives — they
refuse loudly (unsupported quality context mode 5; this archive needs a newer fqxv), never misread. Short-read archives are byte-identical, and
--max(level 9) pins the serial smallest-archive layout, also
byte-identical — chunking is a default-level trade only, and
Params::quality_chunks(0 auto / 1 serial / K ≥ 2) opts out or tunes it. - Nanopore default block budget dropped 256 → 64 MiB of raw sequence
(#275): the benchmark ONT file cuts 5 blocks instead of 2, unlocking
block-parallel decode (measured 3.2× at 8 threads) for +1.1% archive size.
--maxand explicit--block-readskeep maximal blocks. Combined with
chunked quality, the flat ~11 MB/s ONT decode becomes 8.3 → 67 MB/s over
1→32 threads (docs/decode-scaling.md). - Bumped the cargo minor/patch dependency group (#268).
v0.6.2
Patch: no public Rust API changed and the on-disk format is untouched, so every
existing archive reads and writes exactly as before. The changes are additive — a
second unit on the run summary, and type information plus a version attribute on
the Python bindings.
Added
- The Python bindings ship type information. A
py.typedmarker and a
hand-written_fqxv.pyistub give downstream mypy /ty/ IDE users full types
for every class, property, and function instead ofAny, andfqxv.__version__
now reports the installed version. CI type-checks the stubs against the binding
surface (withty) so they cannot silently drift.
Changed
- Run summaries show the binary size beside the SI size — e.g.
archive 211.96 MB (202.14 MiB). fqxv reports decimal SI (MB = 10⁶) while
ls -lh/du -hreport binary units labelledM(MiB = 2²⁰); showing both
reconciles the reported figure with a file listing at a glance. The decimal
number andMBlabel are unchanged. - Bumped
noodles-bgzf0.49 → 0.51.
v0.6.1
Patch rather than minor: no public API changed (every helper involved is
crate-private), no new flag or entry point was added, and the on-disk format is
untouched — group sizes 3 and 4 were always valid and always reachable via
--interleaved, so this only changes when they are chosen. Archives already on
disk are unaffected, and both cross-release compatibility directions are
unchanged.
Fixed
- Single-cell interleaving is no longer archived as paired. Layout detection
only ever compared records pairwise, so a 4-member spot whose members share one
read name — 10x R1/R2/I1/I2 asfasterq-dumpandsrachaemit it, where the
members differ only inlength=— looked like a mate pair at every offset and
was recorded as paired.fqxv inforeported twice the real spot count, and
--splitwrote two mate files that each interleaved two different slots at two
different read lengths, with no warning. Detection now cuts the peeked records
into runs of consecutive same-spot names and infers the common run length (up to
4), so 3- and 4-member layouts are detected and restored to one file per slot.
Archives written before this are unaffected: the stream itself was always
byte-lossless, only the recorded grouping was wrong. - Interleaved SRA deflines are detected as paired.
@RUN.5 5 length=150repeats
the spot name on both mates, and its leading digit was read as a Casava mate
number — so both mates looked like member 5, and since members must be distinct
that vetoed the pairing and the stream was archived single-end, leaving--split
unable to restore the mate files. A leading description digit now counts as a mate
field only when followed by:(1:N:0:…), which is what distinguishes it from a
numeric spot name. This is the defline the documented
sracha get -Z … | fqxv compress -pipeline produces. - Out-of-step inputs warn instead of silently mis-pairing. With
G > 1member
assignment is positional, so a stream missing one member mid-file shifts every
later read into the wrong mate file while the read count stays a clean multiple of
G— invisible to the spot-multiple check and to every CRC. Compression now
compares the members' names within each spot and warns with the count and the
first offending spot. This is the shapefasterq-dump --split-filesproduces for a
spot whoseREAD_LENslot is 0: it writes no record at all for the empty slot. - Two FASTQ shapes SRA conversion produces now have container-level tests, having
been covered only at the codec level: a zero-length read inside aG > 1spot
surviving--splitin its own slot, and.plus the full IUPAC ambiguity set and
lowercase bases round-tripping (SRA's 4na→text map is.ACMGRSVTWYHKDBN, so
converted FASTQ carries more thanN).
v0.6.0
Added
- Golden archive fixtures pin the on-disk format.
crates/fqxv/tests/fixtures/
now holds one small archive per on-disk layout — plain/order-k, grouped (with the
header extension record), the whole-file reorder layout, and the long-read overlap
codec — plus a manifest of what each must decode to. Until now every test in the
workspace compared a build against itself: the round-trips, the proptests, the
fuzz targets, the thread-determinism byte-comparisons, and the accession corpus
(which recompresses from FASTQ each run) are all invariant to a change that moves
the encoder and decoder together, which is exactly the change that breaks archives
already written. Regenerate with
cargo run --release -p fqxv --example make_fixtures -- crates/fqxv/tests/fixtures,
but note that regenerating is almost always the wrong response to a failure. - CI checks cross-release compatibility. A new
compatjob downloads the
previous release's CLI, compresses with it and decompresses with the PR build
(a hard failure if that breaks), then does the reverse and requires the old binary
to either reproduce the bytes exactly or refuse loudly — a silent success with
different output fails the job. - The changelog is on the documentation site. Included verbatim from this file,
so the two cannot drift. - A documented format evolution policy (
docs/design/container.md): which
mechanism a given change belongs in, the rule that any change to the footer's
shape or the block payload's stream layout must be gated by arequired_features
bit, and the project's intent on major bumps (a last resort; a format-1 read path
is retained if one ever happens).
Changed
-
Three implicit parts of the layout now fail closed. Each was a structural
convention living in a constant rather than on disk, and each would accept an
archive from a future writer and get it wrong:- A block payload carrying a fourth stream decoded its first three and dropped
the rest with no error at all — and the per-block content digests, which cover
only the streams that were read, still matched, so the archive looked healthy.
Trailing bytes after the three streams are now rejected. - A footer written at a different per-group stride (a future fourth stream
location) passed its CRC and was then walked at the wrong stride, producing an
index that could satisfy every range check:fqxv inforeported a 1,000-read
archive as 46,048 reads. The footer body length must now match its row-group
count exactly. - Unknown header flag bits (6 and 7) were ignored, so a future flag whose
meaning is purely semantic would decode under the old interpretation with every
CRC and digest still matching. Unknown bits are now refused with a new
Error::UnsupportedFlags.
All three are backward compatible: no archive any release has written trips them.
- A block payload carrying a fourth stream decoded its first three and dropped
-
Original per-slot member labels are recorded and restored. Member identity in
the container is positional, sodecompress_splitcould only number its outputs
_R1.._RG— compressing a run's_2.fastq+_4.fastqand restoring them
renamed them_R1/_R2, and a 10xI1came back as_R3. The CLI now derives
each member's slot token from the input file names (all-or-nothing: every input
must yield a distinct token, or nothing is recorded) and stores them in the
header.--mate-style auto(the new default) restores the stored labels, falling
back to_R1,_R2,…; explicitr/numstay positional. This is the first
user of the header extension region: a non-critical TLV record (tag0x01), so
a reader that predates the tag skips it and decodes byte-identical reads under
positional names. The format version deliberately stays at 1.0 — the record is
skippable, and nothing about decoding depends on it. Archives compressed without
labels still write an empty extension region.
Fixed
--estimate's help text described the wrong mechanism. It claimed to code the
leading reads with the real codecs; it measures the sample's empirical entropy and
projects from that, which is why it is so much faster than a real run. The
documentation was right and the shipped--helpstring was not. A documentation
audit against the actual CLI, bindings, and codecs corrected this and a good deal
more — thefqxv infosample output was internally impossible,-vwas
documented as debug when it is info, batch mode over several archives or a
directory was undocumented, and the container layout diagram omitted the
whole-file reference frame that sits between the header and the first block.fqxv.Index.groups()no longer panics on an allocation failure; it raises.- Unequal input read counts are rejected in both directions.
compress_multi
treated member 0's EOF as a clean end of input, so a short member 0 ended the
archive early and silently dropped the surplus reads from the longer members —
yielding a well-formed, CRC-valid archive that was quietly missing data and still
reported success (with an inflated ratio, since it is computed against the full
input size). All three interleaving sites now confirm members 1..G are also spent.
The reverse direction already errored. - A trailing partial spot is rejected in the reorder layout too. The plain
layout refuses an interleaved stream whose record count is not a multiple of the
group size, but the reorder layout only checked it at one call site, which
compress_autobypasses — so--order any/--maxon a mate-named stream with
an odd record count recordedgroup_size = 2over a stream that is not
spot-aligned and split it into mismatched files. The check moved into
encode_reordered, the choke point every reorder entry point passes through. - The streaming parser validates the
+separator line.read_raw_record
compared only the sequence and quality lengths, so a header line followed by EOF
("@name\n") parsed as a zero-length record — silently repairing a truncated file
into a valid archive — and a garbage third line was consumed and discarded. Both
were reachable only through the multi-input path (a single-file compress rejects
them at the noodles-based peek). The+line's content is still dropped;+
normalization is unchanged. - A non-empty header extension region no longer corrupts the footer.
FooterIndex::newseeded block offsets from the extension-empty header length and
the recovery scan started there, so any extension record shifted every recorded
offset and failed the footer CRC. Both now use the header's actual length. The
region had never been exercised before the member-label record.
v0.5.2
Added
fqxv.estimate()accepts paired-end (and other grouped) inputs. Pass a list
or tuple of sources — paired mates, 10x R1/R2/I1/I2, sharded lanes — and each is
sampled with a split of thesample_readscap, with the per-stream sizes summed
into one aggregateEstimate, matchingfqxv compress … --estimatebyte for
byte. Single-source behavior is unchanged.Estimatealso gains aplatform
field ("illumina","nanopore","pacbio","mgi","unknown").- Cross-platform guard on grouped estimates. A group whose inputs resolve to
two different known platforms is now a hard error in both the Python API and
the CLI's--estimate, instead of silently summing two differently-calibrated
numbers. AnUnknowninput (SRA-renamed, no content signal) stays compatible
with anything, so real SRA data never trips a false positive.
Changed
- New
fqxv-aligncrate. The two edit-distance alignment implementations —
banded Needleman-Wunsch (AVX2 + scalar) and the wavefront aligner — moved out of
fqxv-lroverlapinto a zero-dependency leaf crate, so an external consumer can
reach them without pulling in the whole long-read codec cone. The move is a pure
rename:fqxv-lroverlapre-exports all seven public names at their existing
paths, and its public surface is unchanged.fqxv-lroverlapis now
#![forbid(unsafe_code)], since the AVX2 backend that motivated the weaker
denyhas left the crate. - Dependency bumps:
noodles-bgzf0.48 → 0.49,clap4.6.2 → 4.6.3,
xxhash-rust0.8.17 → 0.8.18,serde_json1.0.150 → 1.0.151.
v0.5.1
Performance
fqxv compress --estimateis now 15–60× faster — always subsecond on a
block-sized sample. It predicts the archive from the sample's measured entropy
(what the real entropy coders converge to) instead of coding it: static order-k
sequence entropy, a k-mer duplication sketch that captures long-read cross-read
redundancy, order-1 quality entropy, and the real tokenizer for names — each
scaled by a small per-platform calibration factor. The measurement passes run in
parallel across read chunks, and the sample is capped by read count for short
reads and by one block of bases for long reads (so the sketch sees a real block's
coverage). Predicted archive size stays within ~1% of the previous coding-based
estimate across Illumina, PacBio HiFi, and Nanopore.
v0.5.0
Added
fqxv decompress -reads from stdin, streaming the archive straight into the
decoder (it only reads forward and stops at the terminator frame). This is how you
read a remote archive — pipe a transfer tool in, so all the auth, retries, and
resume stay in the tool built for it and fqxv gains no HTTP dependency:
aws s3 cp s3://bucket/reads.fqxv - | fqxv decompress - -Z | bwa mem ref.fa -
(orcurlon a presigned URL,gsutil cat, …). A truncated stream still fails
(premature EOF) rather than yielding a short file.--recoverand--splitneed
a seekable file (a stream can't be rewound) and refuse stdin with a clear message.- Python: read archives over the network. The
fqxvwheel gains a
dependency-freefqxv.remotemodule (standard-libraryurllib). The streaming
entry points —fqxv.open,decompress_to_*— now accept any file-like
object, so aboto3/urllib/httpxresponse streams straight in
(fqxv.open(s3.get_object(...)["Body"]));fqxv.remote.stream(url)/
download(url, dest)wrap that for a URL.fqxv.remote.RemoteArchiveadds
column projection over HTTP byte-range requests: fetch just the footer index
from the archive tail, then only the names (~1% of the file) or the sequence,
CRC-verified per stream. It rests on IO-free primitives —
parse_index_suffix, per-stream ranges/CRCs onIndex, and
decode_{names,sequences,qualities}_bytes— that a custom async client can drive
directly for concurrent range fetches.
Performance
- Reorder compress: SIMD merge scan, malloc-free placement, incremental
cast_vote. Three byte-identical, zero-ratio-cost speedups on the hot
short-read reorder assembly/merge phases (--order any/shuffle/--max).
(1) The overlap-merge successor scan now compares 16 bytes at a time
(_mm_cmpeq_epi8+ movemask popcount, budget checked per block) behind a
runtimeis_x86_feature_detected!("sse2")dispatch with a byte-identical
scalar fallback (no raised global baseline) — the count matches the scalar
loop exactly when within budget and exceeds it exactly when the scalar loop
would break, sobest_keyand the archive are unchanged. (2)place_on_contig
and the rescue assembler'stry_place/placeno longer allocate a
Vec<usize>of mismatch positions per candidate: candidates are scored with a
count-only pass and only the winner's positions are materialized (into a
caller-reused scratch buffer on the clustered path). (3)cast_voteupdates
the per-column plurality in O(1) by comparing the just-incremented base against
the current winner (lowest-index tie rule preserved) instead ofmax_by_key
over all four counts. About 2% faster single-thread (54.87 s → 53.83 s,
1.02× ± 0.00, on 1M-read NovaSeq--order shufflecompress); within noise at
--threads 16(1.00× ± 0.04), where the wall clock is dominated by the serial
assembly prelude. Byte-identical archives on the validation datasets
(NovaSeq/GAIIx/MiSeq/MGI, both--order shuffleand--max), and
thread-deterministic (--threads 1==--threads 16). (#219) - Skip the quality-quantizer probe on Nanopore. The per-block quantizer trial
(MODE_SEQ_BINMIX_Q) only wins on skewed PacBio HiFi/Revio quality; on the flatter
Nanopore distribution it never does, yet the bounded prefix probe still ran and cost
~5% of ONT compress for 0 bytes saved.encode_seqnow takes atry_quantizerflag
and the container passesfalseon Nanopore, skipping the histogram build and the
probe entirely. Output is byte-identical (the kept stream was the baseline either
way); only ONT compress speed changes. HiFi/Revio still trial and keep the quantizer.
Internal
-
Optimize the workspace codec crates in dev/test builds. The
[profile.dev]
package."*"override optimized external dependencies but not workspace members, so
the codec crates compiled at opt-level 0 for tests — and the test suite runs real
compression, so the long-read round-trip tests took 30–45 s each. Per-package
opt-level = 2overrides for the compute crates cut the CI test suite roughly 4–5×
with no behavior change (SIMD is runtime-detected; determinism holds across opt-levels).
The CLI stays at opt-level 0 to keep the edit-compile loop fast. -
Byte-identical speed in the long-read overlap search (chaining). The minimizer
overlap search (fqxv-lroverlap) is the top self-cost of Nanopore compress, and a
overlap search (fqxv-lroverlap) is the top self-cost of Nanopore compress, and a
query's per-anchor chaining dominates it (a fresh profile putChainer::chainat
~23% andfind_overlapsat ~14% of ONT compress). Three output-preserving changes
strip wasted work without moving a single archive byte:- The chainer's redundant re-sort is dropped.
find_overlapsalready sorts a
query's whole anchor set by(target, strand, tpos, qpos), so each
(target, strand)group it hands the chainer arrives already in(tpos, qpos)
order — exactly what the chainer's ownsort_unstablewould produce. A new
Chainer::chain_presortedentry point skips that re-sort (it still dedups),
saving hundreds of sorts per read at high coverage. - The chaining DP's scratch buffers are reused. The per-anchor score,
predecessor, order, and used arrays — plus the per-chain path — were
heap-allocated on every group (hundreds of small allocations per read). They are
now thread-local buffers, cleared and resized per group, never read stale — so
the chain set stays a pure function of the input. - The hot anchor-bucket sort is a radix sort. Each anchor's group-and-position
key(target, strand, tpos, qpos)packs losslessly into au128whose
ascending order is exactly the tuple order, so the per-query sort (~7% of ONT
compress) is now a deterministic LSD radix sort producing the identical total
order rather than a comparison sort.
About 15% faster single-thread ONT compress (ecoli_ont/DRR205413,
559 s → 475 s), where chaining is the largest self-cost. At 16 threads long-read
compress is block-bound — a handful of large blocks gate the wall clock, not the
per-anchor work — so the gain is dataset-shaped: ~noise on ONT (few, large blocks)
but about 10% faster on high-coverage HiFi (ecoli_hifi,--platform pacbio,
254 s → 228 s at 16 threads), where the many smaller blocks keep the cores fed.
Archives are byte-identical on ONT (default and--max) and HiFi, and
thread-count invariant (--threads 1==--threads 16). Proptests pin the radix
order to the comparison sort's and the presorted chain path to the sorting path.
The levers mirror minimap2's own presorted chaining andradix_sort_128x. (#151)
- The chainer's redundant re-sort is dropped.
-
Per-block quality context quantizer, trialled and kept only when it wins
(HiFi/Revio). The long-read quality coder used to build its recent-quality
context with one fixed quantization (q1>>1/q2>>3/q3>>4), which merges
adjacent Phred values — costly on HiFi/Revio, where quality is packed at the top
of the scale and neighbouring high-Q values need to be told apart. The encoder now
also builds a per-block quantizer from the block's quality histogram (the fqzcomp
qtab/ CoLoRd platform-quantizer idea): full context resolution where quality
actually varies, equal-population folding elsewhere. Both quantizers code the
block and the smaller is kept, with the choice recorded in a self-describing
header mode byte (MODE_SEQ_BINMIX_Q) and the small table transmitted so decode is
unambiguous — so a block can only match or shrink (never-worse by construction).
Quality-stream savings, measured against a clean baseline: PacBio HiFi (ecoli,
Sequel II) −1.17%, Revio amplicon −3.25%, Revio WGS −0.25%; Nanopore
0% and short reads 0% (byte-identical) — neither is touched. A bounded
prefix probe decides whether the second full encode is worth coding, so the trial
costs ≈ +5% compress on Nanopore (where it never wins) and is paid in full only on
the skewed long-read data where it does. Lossless and thread-deterministic;
archives round-trip byte-for-byte and are identical regardless of thread count. -
Anchor-restricted long-read tile coding (CoLoRd-style). The multi-reference
ONT tiler used to re-align each tile with one banded DP over the whole
read × referencewindow, re-deriving the exact-match stretches its own minimizer
chain had already proven identical. It now walks that chain, emits each anchor as a
freeMatchcopy-run, and runs the aligner only on the short inter-anchor gaps and
the two flanks — so alignment work scales with divergence, not read length. About
2.5× faster single-thread ONT tiler compress, with an aggregate ratio
improvement across the ONT corpus (most accessions smaller — e.g. DRR205413 −2.9%,
DRR351396 −2.8%, DRR424350 −2.6%; worst real-data case ≈ +0.3%). Lossless and
thread-deterministic; the edit-op stream keeps the same semantics, so the decoder
is unchanged and archives still round-trip byte-for-byte. A tile whose chain is not
recovered falls back to the whole-window DP;FQXV_TILE_NO_ANCHORGAP=1forces the
DP everywhere for A/B. (#226) -
Anchor-restricted coding for the shared-reference consensus codec (HiFi). The
same lever as #226, now on the non-tiler path the high-coverage HiFi/PacBio blocks
take: each read used to be re-aligned against its consensus with one banded DP over
the wholeread × consensuswindow (align_banded, the dominant HiFi compress
self-cost). It now recovers the read↔consensus exact-k-mer chain, emits each anchor
as a freeMatchcopy-run, and aligns only the short inter-anchor gaps and two
flanks — so alignment work scales with the read's divergence from its consensus,
not its length. On low-error HiFi the read shares nearly all of it...
v0.4.0
Added
-
Best-of-N reference selection takes the ONT tiler to CoLoRd parity — the
multi-reference tiling codec (SEQ_METHOD_TILE) now weighs several earlier-read
references per tile and keeps the cheapest edit script, instead of blindly taking
the furthest-reaching neighbour. At ONT coverage many earlier reads span the same
region with independent error patterns, so the lowest-cost reference agrees with
the query at more positions — the way CoLoRd picks its anchor, applied per tile.
With a wider alignment band on top this is the dominant ONT sequence-ratio lever:
on theecoli_ontarchive the sequence stream drops from 43.3 MB to 34.5 MB
(1.150 → 0.915 bits/base), matching CoLoRd's ~34 MB, and the whole archive goes
2.92× → 3.05× — losslessly (--verify). It is encoder-only (the tile block
self-describes, so the decoder is unchanged and single-reference coding is
byte-for-byte identical to before) and compute-heavy (~linear in the fan-out on
the tiler's already-vectorised alignment), so it is gated to the top effort
levels: the default keeps the single-reference cover and today's ONT speed,
-l7/-l8enable best-of-2/-4, and--maxreaches the parity operating point.
Nanopore long reads only.FQXV_TILE_BAND/FQXV_TILE_REFSoverride the band
and fan-out for rebuild-free A/B measurement. -
Raw-LZMA sequence codec for ordinary-coverage long reads — a new per-block
sequence method (SEQ_METHOD_LZMA) codes the block's ASCII bases with a
large-window LZ, kept only when it beats the existing candidates. At ordinary
long-read coverage — a real genome, not 300× of one organism — overlapping HiFi
reads share long exact substrings that neither the within-read order-k model
nor the consensus-edit overlap codec captures, but a large-window LZ finds
directly. On the full PacBio Revio WGS run the archive goes from 9.72× to
13.35× — from second-worst in the lossless field, below plainzstd19/xz9,
to above both — losslessly. The overlap codec still wins on
high-coverage/amplicon data, and the per-blockkeep_smallergate picks the
right one automatically.FQXV_SEQ_NO_LZMAdisables the candidate for A/B
measurement.
Changed
-
--versionreports git provenance on development builds. A binary built
from a clean checkout of the release tag still prints justfqxv 0.4.0;
anything else — commits past the tag, a dirty tree, an untagged branch —
appends the git description (fqxv 0.4.0 (v0.4.0-7-gab12cd34-dirty)), so a
bug report identifies the exact build. Source trees with no git (a crates.io
tarball) print the plain version. -
Real progress bars for paired compress and decompress. Both paths showed a
bare indeterminate spinner for the whole run — minutes of no information on a
paired sample at a few cores. Compress picked its indicator on input count, so
paired mates always fell through to the spinner even though the summed input
length the bar needs was already computed; interleaving reads every input in
lockstep, so one shared counter against that sum drives the same bar the
single-input path already used. Decompress had no counter at all and now reports
reads finished against the footer's read count (already read up front for the
truncation check) — counting output records, not archive bytes consumed, which in
--threads-sized batches runs ahead of the work and would pin at 100% while the
last batch is still decoding. Both bars now carry an ETA. The reorder and stdin
carve-outs still degrade to an indeterminate readout rather than report a
percentage that would stall or lie, and--quietand non-TTY output are
unchanged (#189).
Performance
-
Lower peak memory on Nanopore compression. The minimizer-occurrence record
(Occ) that dominates the long-read overlap index is packed from 24 bytes to 16
— the read index and the strand flag now share oneu32(read in the low 31
bits, strand in the high bit) instead of au32beside aboolthat padded the
struct out to a third 8-byte word. That is a one-third cut of the index's largest
allocation, on the order of a gigabyte on a large ONT block, so it directly eases
the Nanopore memory pressure. Output is byte-identical — a hand-writtenOrd
reproduces the previous(hash, read, pos, strand)sort order exactly — and that
was verified both by the determinism round-trips and by an archive-level diff
against the previous build. -
Nanopore compression is ~2.3x faster — the shared whole-file reference layout
(#168) is now skipped on high-error Nanopore, where it never pays off. It helps
low-error PacBio HiFi (a clean consensus stored once), but on noisy ONT the
reference frame always costs more than it saves, so the whole-file gate (#184)
rejects it every time — after building the whole-file reference and coding every
block against it, a second full long-read assembly per block that profiling put at
~45% of ONT compress CPU, all discarded. Skipping it on Nanopore goes straight to
the plain layout the gate would fall back to regardless, so the output is
byte-identical — measured 7:20 → 3:11 on a 600 MB E. coli ONT file. HiFi is
unaffected. Mirrors the Nanopore LZMA gate above. -
PacBio HiFi is smaller and ~2x faster to compress — the raw-sequence LZMA
codec (SEQ_METHOD_LZMA, #197) now seeds its match finder on 12 bytes
instead of 4. On raw ASCII a 4-byte seed keys only ~256 distinct DNA 4-grams, so
the hash chains collided catastrophically and the depth cap truncated them —
which both cost time (millions of pointer-chases through a capped chain) and
lost matches beyond the cap. A 12-base seed keys ~16.7M grams, so a 12-mer
recurs only ~a dozen times: chains are short, the cap never binds, and the
finder sees every candidate. On a 40k-read Revio WGS subset the sequence stream
drops 0.791 → 0.638 b/base and the whole archive 29.4 → 25.2 MB
(−14%), while compress time falls ~2x (329s → 166s, below even the
original 4-byte path). Validated lossless and on a second HiFi dataset. The
match finder is tuned per call site, so the 2-bit packed-reference path keeps
its 4-byte seed and byte-identical output; decode is seed-agnostic, so no format
change. The match-extension inner loop is also AVX2-vectorized (CPU-feature
gated, byte-identical, ~8% on top;FQXV_LZMA_NO_SIMDforces scalar). -
Nanopore compression is ~5.5x faster — the raw-LZMA sequence candidate
(SEQ_METHOD_LZMA) is now skipped on high-error Nanopore, where it cannot win.
LZMA pays off only where reads share long exact substrings — ordinary-coverage
low-error PacBio HiFi (0.60 vs the overlap codec's 1.39 b/base); on Nanopore a
base error every ~10 bases chops those matches short, so it reliably loses to
the overlap/order-k codecs (2.2 vs 1.79 b/base) while costing the most encode
time of any candidate. Because it runs on the serial block-0 probe as well as
the fan-out blocks, coding-and-discarding it dominated the ONT wall-clock: on a
600 MB E. coli ONT run compress drops 1362s → 246s with byte-identical
output (LZMA was ~82% of the time and changed nothing). The gate keys off the
detected platform, so PacBio still codes LZMA and keeps its win, and any
platform where the loss can't be ruled out up front (Unknown) still codes it —
the never-worse ratio guarantee is unchanged.keep_smallercontinues to floor
every block regardless. -
Long-read compression is ~40% faster where the shared reference wins
outright. Making the whole-file gate exact meant coding both layouts for
every block, which costs a second long-read assembly per block — on PacBio
HiFi that roughly doubled compress time (225s → 461s onecoli_hifi) to
evaluate a candidate that never had a chance. The encoder now probes only the
first block both ways and, when the reference wins by a wide enough margin
to cover the reference frame three times over, skips the plain candidate for
the remaining blocks.ecoli_hifidrops 461s → 275s with byte-identical
output;ecoli_ont, where the plain layout genuinely wins, sees a negative
margin and still runs the exact gate, so the ONT ratio is unchanged. The probe
is integer arithmetic over block 0, so the decision stays thread-count
invariant, and a shortcut block is still floored at order-k.
Fixed
-
Higher effort no longer produces a larger archive. On a low-redundancy
library (BGISEQ-500)--maxand-l9compressed worse than the default,
because two effort knobs were applied unconditionally even where they cost more
than they saved (#196). Both are now never-worse: (1) the level-8+ hashed
high-order sequence tier is coded alongside plain order-k and the smaller kept,
so enabling it can only help; (2) single-end--order any/--maxcodes the
clustered and plain layouts and keeps the smaller, so reordering is used only
when it pays — the same "code both, keep the smaller" rule the long-read shared
reference (#192) and the overlap-vs-order-k choice already use. On the BGISEQ
case--maxdrops from a 340 KB regression to matching the default; on a
reorder-favourable NovaSeq set--maxkeeps its 2.4x win. Grouped (paired /
10x) reorder is unchanged: it pays the permutation for spot reconstruction
regardless, so the tradeoff differs. -
The long-read shared-reference gate now measures against the layout it falls
back to. Adoption of the whole-file reference was gated on beating the plain
order-k total, but the fallback path keeps the smaller of the per-block
overlap codec and order-k — a stronger result than the bar being tested. A
reference that lost to the per-block overlap codec could therefore still be
adopted. Both layouts are now coded in the first pass and the smaller wins, so
the never-worse property holds against the real alte...
v0.3.0
Added
-
Sequence-conditioned, context-mixed quality for long reads — the quality
coder now conditions each score on the read's bases and recent qualities instead
of read position (which carries no signal on long reads). It mixes several
context models of increasing richness (coarse/mid/rich) with adaptive,
confidence-gated weights — a logistic mixer — because a per-block adaptive model
can't exploit a richer single context but can blend a well-trained coarse one
with a sparse rich one. On PacBio HiFi this takes the quality stream (the dominant
share of a HiFi archive) below CoLoRd, lossless. The mode is chosen automatically
by mean read length and recorded in a self-describing header byte, so short-read
archives are byte-identical to before; long-read archives decode the sequence
first and feed it to the quality decoder. The mixer is fixed-point/integer
throughout, so archives are bit-identical across platforms. NewfqzcompAPI:
encode_seq/decode_seq/needs_sequence; new random-access projection helpers
decode_quality_with_seq/quality_needs_sequence. -
Long-read (ONT / PacBio) support — platform-aware compression for
Nanopore and PacBio reads:--platform illumina|nanopore|pacbio(with header
auto-detection), long-read quality binning (--quality-bin ont|hifi, matching
CoLoRd's cutpoints), and a dedicated long-read overlap sequence codec that
drives its minimizer sketch from the detected platform. -
Long-read overlap sequence codec (
fqxv-lroverlap) — a new cross-read
overlap codec (minimizers → overlaps → layout → consensus → per-read banded
edit script → rANS) wired into the container as the sequence path for long-read
blocks, auto-selected and kept only when it beats order-k. It reaches CoLoRd
parity on ONT/HiFi (e.g. ~0.653 → ~0.067 bits/base at depth) and codes its edit
streams (substitutions, ops, insertions) with per-stream context models rather
than a flat order-0 code — worth a further −4.3% on the ONT sequence stream,
losslessly. A WFA aligner (HiFi) and an AVX2 anti-diagonal aligner accelerate
encoding, both byte-identical to the scalar reference. -
Closed-syncmer seeding for ONT — the overlap codec seeded from window
minimizers, which a base error in any k-mer of the window can deselect; at
Nanopore's ~10% error that is the dominant loss of shared anchors. ONT now seeds
with closed syncmers (a k-mer is an anchor when its minimals-mer sits at its
first or last position), so selection depends only on the k-mer's own bases and
an intact shared k-mer is co-selected regardless of neighbouring errors. Density
is2/(w+1)as before, sok, anchor count and specificity are unchanged — only
conservation improves: ONT overlap-coded sequence goes 1.631 → 1.559 bits/base
against an order-k baseline of 1.808, widening with per-block coverage. Seeding
is encode-only, so archives stay decodable by any reader. PacBio keeps window
minimizers, which are already near-optimal below ~1% error. -
Content-based platform detection — platform was detected from read-name
grammar alone, so SRA-reformatted runs (bareSRR…headers) recordedunknown
and were handed the Nanopore sketch. When names carry no platform signal, fqxv
now classifies long-read runs by mean per-base quality, which separates the
platforms cleanly (measured across 24 corpus accessions with ENA ground truth:
Nanopore 6.9–23.5, PacBio 36.9–84.5). Ambiguous data staysunknown, so a wrong
platform is still never recorded. On a real HiFi run this restores the
low-divergence WFA path and cuts encode time ~31%. -
Shared whole-file reference for long reads — the overlap codec reaches
CoLoRd parity within a block, but the container re-assembled and re-stored the
same consensus reference in every 256 MiB block. It now assembles one consensus
over the whole file and stores it once in a framed region between the header
and the first block, coding every block's reads against that frozen frame — so
the genome is stored once, not once per block. Auto-selected for long-read input,
behind a whole-file never-worse gate against order-k, and gated on the
GLOBAL_REFERENCEfeature bit (which is set in the plain layout now as well as
the reorder layout, so a reader without it refuses rather than misreads). On
PacBio HiFi this drops the sequence stream ~0.102 → ~0.084 bits/base at two
blocks, widening with block count. Newfqxv-lroverlapAPI:Reference,
build_reference,encode_against/decode_against. -
Python bindings (
fqxvvia PyO3 / maturin) — read-only access from
Python: a streaming record iterator and column projection (fetch just names,
sequence, or quality) over an existing archive. -
compress --verify— an opt-in read-after-write check that round-trips the
freshly written archive back to the original records before exiting, backed by a
publicverify_roundtripin the library. -
Block sync markers + footer-independent recovery — each block carries a
sync marker so a reader can resynchronize and recover blocks even when the
footer index is missing or truncated. -
compress -f/--force— compression now refuses to overwrite an existing
output unless--forceis given. -
Remote / parallel column projection — the footer row-group index now
records, per group, a(offset, len, crc32c)triple for each of the three
coded streams (names, sequence, quality). A client can fetch the archive tail,
parse the index, and issue a single range request for just one stream — read
names are <1% of the archive, so an ID-only client fetches ~100× less than
before, and a sequence-only client (k-mer screening, classification) skips the
quality stream. The joint block content digest can't verify a single fetched
stream, so each stream carries its own CRC-32C. -
Random-access API — a public, IO-free
Index(parse from a seekable
reader or a fetched suffix buffer viaIndex::from_suffix),Index::byte_ranges
to turn(groups, stream)into byte ranges toGET,Index::verify_stream,
and per-stream (decode_names/decode_sequence/decode_quality) and
whole-block (decode_block_contents) decoders. The caller drives fetching
(localFile,object_store, async HTTP, …). -
compress --block-reads N— set the reads-per-row-group directly,
decoupling random-access granularity from the--leveleffort knob. Smaller
groups give finer remote access and more parallelism at some ratio cost. -
Per-stream content digests — the plain block payload now carries three
xxh3-64 digests (names, sequence, quality) in place of the single joint digest.
A post-decode mismatch (a codec round-tripping CRC-valid bytes into
wrong-but-in-bounds output) now names the offending stream instead of only the
block. Each digest still folds inn_readsand its stream's per-read lengths,
so boundary pinning is unchanged; cost is 16 extra bytes per block. -
Crash-safe compress output — compression writes to a sibling temp file and
atomically renames it into place only once the whole archive (header, blocks,
and footer trailer) is on disk. Interrupting a stream mid-run (Ctrl-C on
sracha get -Z | fqxv compress -, say) or hitting any error no longer leaves a
corrupt, footer-less.fqxvat the destination: the partial temp is removed on
a?bail and by a SIGINT/SIGTERM/SIGHUP handler, and the destination path only
ever holds a complete archive. -
Live compress progress — the compress indicator now reports how much data
has been processed and the rate. With a known input size (a file) it renders a
percentage bar; for a stdin stream of unknown length it shows a bytes + rate
readout, so a pause waiting on an upstream producer reads as0 Brather than a
hang.
Changed
- The on-disk format is stable at 1.0. Versioning moved from a single
monotonicFORMAT_VERSIONinteger toFORMAT_MAJOR.FORMAT_MINOR: a reader
refuses a differing major and tolerates a newer minor, and additive features are
gated behind required-feature bits so a reader that predates a feature refuses
the archive outright rather than misreading it. That contract is now a stability
guarantee — archives written by a 1.x release stay readable by later ones, and a
major bump would be announced as a breaking change. The 1.0 format carries the
long-read overlap codec, the extended per-stream footer index, and the
per-stream block digests below. platformhas its own header byte — the platform tag no longer shares bits
with the flags byte (a prior collision made Illumina--order anyarchives
undecodable).inspectnow sums per-stream sizes straight from the footer index instead
of seeking to each block header — one footer read is the whole metadata cost.fqxv-dnaprimitives crate — the 2-bit ACGT lookup and reverse-complement
helpers are extracted into a shared leaf crate, and thefqxv-reordermonolith
is split into focused modules.- Default log verbosity is
warn— routine per-runinfodiagnostics now
require-v, so they no longer interleave with and smear the live compress
indicator on a shared stderr (-vv= debug,-vvv= trace with targets). --verifyverifies before publishing — the read-after-write check now runs
against the temp file, so a failed verification leaves the unverified output off
the destination path (kept aside for inspection) and exits non-zero, rather than
leaving a suspect archive in place.--helplayout — the flags most runs need are listed first and the
rarely-touched knobs are grouped under anAdvancedheading, so the default
help is short enough to read.-hprints a one-line summary per flag and
--helpexpands the detail (when a knob changes the data or the guarantees,
the long form says so).- Rust edition 2024 — the ...
v0.2.0
Added
compress --estimate— predict the compression ratio and archive size
for a FASTQ without writing an archive.--estimate tsvemits machine-readable
output, andinfo/verifynow accept multiple files or directories for batch
reporting.- Reference sequence coder for
--max/reorder — a SPRING-style codec that
2-bit-packs the assembled global reference and entropy-codes it with a
clean-room LZMA, adopted only when it beats the raw representation. --order shufflerenumber mode — a true SPRING-style read renumbering that
discards the input permutation, reaching SPRING-competitive ratios on datasets
where read order carries no information.- CLI run feedback — a TTY-aware progress spinner and human-readable
compress/decompress run summaries (suppressed under--quietand when stderr
is not a terminal).
Changed
- Quality coding — fqzcomp now models quality over the symbol alphabet that
actually occurs in the file rather than a fixed0..QMAXrange, shrinking the
quality stream on data with sparse quality alphabets. Byte-identical
round-trips are preserved.
Performance
- Reorder merge — the overlap-merge k-mer index now uses a rolling hash with
sharding (~10% faster--max, byte-identical output), and the
merge_referencevote-scatter is parallelized.
Documentation
- VHS-rendered terminal demos on the docs site, refreshed benchmarks (including
anfqxv-shufflerow), and clearer positioning.