v0.3.0
Added
-
Sequence-conditioned, context-mixed quality for long reads — the quality
coder now conditions each score on the read's bases and recent qualities instead
of read position (which carries no signal on long reads). It mixes several
context models of increasing richness (coarse/mid/rich) with adaptive,
confidence-gated weights — a logistic mixer — because a per-block adaptive model
can't exploit a richer single context but can blend a well-trained coarse one
with a sparse rich one. On PacBio HiFi this takes the quality stream (the dominant
share of a HiFi archive) below CoLoRd, lossless. The mode is chosen automatically
by mean read length and recorded in a self-describing header byte, so short-read
archives are byte-identical to before; long-read archives decode the sequence
first and feed it to the quality decoder. The mixer is fixed-point/integer
throughout, so archives are bit-identical across platforms. NewfqzcompAPI:
encode_seq/decode_seq/needs_sequence; new random-access projection helpers
decode_quality_with_seq/quality_needs_sequence. -
Long-read (ONT / PacBio) support — platform-aware compression for
Nanopore and PacBio reads:--platform illumina|nanopore|pacbio(with header
auto-detection), long-read quality binning (--quality-bin ont|hifi, matching
CoLoRd's cutpoints), and a dedicated long-read overlap sequence codec that
drives its minimizer sketch from the detected platform. -
Long-read overlap sequence codec (
fqxv-lroverlap) — a new cross-read
overlap codec (minimizers → overlaps → layout → consensus → per-read banded
edit script → rANS) wired into the container as the sequence path for long-read
blocks, auto-selected and kept only when it beats order-k. It reaches CoLoRd
parity on ONT/HiFi (e.g. ~0.653 → ~0.067 bits/base at depth) and codes its edit
streams (substitutions, ops, insertions) with per-stream context models rather
than a flat order-0 code — worth a further −4.3% on the ONT sequence stream,
losslessly. A WFA aligner (HiFi) and an AVX2 anti-diagonal aligner accelerate
encoding, both byte-identical to the scalar reference. -
Closed-syncmer seeding for ONT — the overlap codec seeded from window
minimizers, which a base error in any k-mer of the window can deselect; at
Nanopore's ~10% error that is the dominant loss of shared anchors. ONT now seeds
with closed syncmers (a k-mer is an anchor when its minimals-mer sits at its
first or last position), so selection depends only on the k-mer's own bases and
an intact shared k-mer is co-selected regardless of neighbouring errors. Density
is2/(w+1)as before, sok, anchor count and specificity are unchanged — only
conservation improves: ONT overlap-coded sequence goes 1.631 → 1.559 bits/base
against an order-k baseline of 1.808, widening with per-block coverage. Seeding
is encode-only, so archives stay decodable by any reader. PacBio keeps window
minimizers, which are already near-optimal below ~1% error. -
Content-based platform detection — platform was detected from read-name
grammar alone, so SRA-reformatted runs (bareSRR…headers) recordedunknown
and were handed the Nanopore sketch. When names carry no platform signal, fqxv
now classifies long-read runs by mean per-base quality, which separates the
platforms cleanly (measured across 24 corpus accessions with ENA ground truth:
Nanopore 6.9–23.5, PacBio 36.9–84.5). Ambiguous data staysunknown, so a wrong
platform is still never recorded. On a real HiFi run this restores the
low-divergence WFA path and cuts encode time ~31%. -
Shared whole-file reference for long reads — the overlap codec reaches
CoLoRd parity within a block, but the container re-assembled and re-stored the
same consensus reference in every 256 MiB block. It now assembles one consensus
over the whole file and stores it once in a framed region between the header
and the first block, coding every block's reads against that frozen frame — so
the genome is stored once, not once per block. Auto-selected for long-read input,
behind a whole-file never-worse gate against order-k, and gated on the
GLOBAL_REFERENCEfeature bit (which is set in the plain layout now as well as
the reorder layout, so a reader without it refuses rather than misreads). On
PacBio HiFi this drops the sequence stream ~0.102 → ~0.084 bits/base at two
blocks, widening with block count. Newfqxv-lroverlapAPI:Reference,
build_reference,encode_against/decode_against. -
Python bindings (
fqxvvia PyO3 / maturin) — read-only access from
Python: a streaming record iterator and column projection (fetch just names,
sequence, or quality) over an existing archive. -
compress --verify— an opt-in read-after-write check that round-trips the
freshly written archive back to the original records before exiting, backed by a
publicverify_roundtripin the library. -
Block sync markers + footer-independent recovery — each block carries a
sync marker so a reader can resynchronize and recover blocks even when the
footer index is missing or truncated. -
compress -f/--force— compression now refuses to overwrite an existing
output unless--forceis given. -
Remote / parallel column projection — the footer row-group index now
records, per group, a(offset, len, crc32c)triple for each of the three
coded streams (names, sequence, quality). A client can fetch the archive tail,
parse the index, and issue a single range request for just one stream — read
names are <1% of the archive, so an ID-only client fetches ~100× less than
before, and a sequence-only client (k-mer screening, classification) skips the
quality stream. The joint block content digest can't verify a single fetched
stream, so each stream carries its own CRC-32C. -
Random-access API — a public, IO-free
Index(parse from a seekable
reader or a fetched suffix buffer viaIndex::from_suffix),Index::byte_ranges
to turn(groups, stream)into byte ranges toGET,Index::verify_stream,
and per-stream (decode_names/decode_sequence/decode_quality) and
whole-block (decode_block_contents) decoders. The caller drives fetching
(localFile,object_store, async HTTP, …). -
compress --block-reads N— set the reads-per-row-group directly,
decoupling random-access granularity from the--leveleffort knob. Smaller
groups give finer remote access and more parallelism at some ratio cost. -
Per-stream content digests — the plain block payload now carries three
xxh3-64 digests (names, sequence, quality) in place of the single joint digest.
A post-decode mismatch (a codec round-tripping CRC-valid bytes into
wrong-but-in-bounds output) now names the offending stream instead of only the
block. Each digest still folds inn_readsand its stream's per-read lengths,
so boundary pinning is unchanged; cost is 16 extra bytes per block. -
Crash-safe compress output — compression writes to a sibling temp file and
atomically renames it into place only once the whole archive (header, blocks,
and footer trailer) is on disk. Interrupting a stream mid-run (Ctrl-C on
sracha get -Z | fqxv compress -, say) or hitting any error no longer leaves a
corrupt, footer-less.fqxvat the destination: the partial temp is removed on
a?bail and by a SIGINT/SIGTERM/SIGHUP handler, and the destination path only
ever holds a complete archive. -
Live compress progress — the compress indicator now reports how much data
has been processed and the rate. With a known input size (a file) it renders a
percentage bar; for a stdin stream of unknown length it shows a bytes + rate
readout, so a pause waiting on an upstream producer reads as0 Brather than a
hang.
Changed
- The on-disk format is stable at 1.0. Versioning moved from a single
monotonicFORMAT_VERSIONinteger toFORMAT_MAJOR.FORMAT_MINOR: a reader
refuses a differing major and tolerates a newer minor, and additive features are
gated behind required-feature bits so a reader that predates a feature refuses
the archive outright rather than misreading it. That contract is now a stability
guarantee — archives written by a 1.x release stay readable by later ones, and a
major bump would be announced as a breaking change. The 1.0 format carries the
long-read overlap codec, the extended per-stream footer index, and the
per-stream block digests below. platformhas its own header byte — the platform tag no longer shares bits
with the flags byte (a prior collision made Illumina--order anyarchives
undecodable).inspectnow sums per-stream sizes straight from the footer index instead
of seeking to each block header — one footer read is the whole metadata cost.fqxv-dnaprimitives crate — the 2-bit ACGT lookup and reverse-complement
helpers are extracted into a shared leaf crate, and thefqxv-reordermonolith
is split into focused modules.- Default log verbosity is
warn— routine per-runinfodiagnostics now
require-v, so they no longer interleave with and smear the live compress
indicator on a shared stderr (-vv= debug,-vvv= trace with targets). --verifyverifies before publishing — the read-after-write check now runs
against the temp file, so a failed verification leaves the unverified output off
the destination path (kept aside for inspection) and exits non-zero, rather than
leaving a suspect archive in place.--helplayout — the flags most runs need are listed first and the
rarely-touched knobs are grouped under anAdvancedheading, so the default
help is short enough to read.-hprints a one-line summary per flag and
--helpexpands the detail (when a knob changes the data or the guarantees,
the long form says so).- Rust edition 2024 — the workspace moved from edition 2021. The MSRV is
unchanged at 1.95, so this is invisible to downstream builds on a supported
toolchain.
Performance
- Single-end compress now streams through a block pipeline instead of
buffering the whole input. - Reorder read storage flattens
cl_readsinto an arena, cutting ~168 MB of
peak memory with byte-identical output. - Quality models are sized to the alphabet actually present in the file
(~17% faster quality coding, ~45 MB less memory).
Fixed
- Silent quality corruption on certain inputs and a
--block-reads
compression-bomb path, both surfaced by CLI stress testing. - fqzcomp no longer rejects legitimately compressible quality (which could
drop data). - Read-name headers are preserved byte-exactly.
- Long-read memory — the overlap codec's overlap/layout/placement and the
banded aligner's DP matrix are now bounded, fixing OOMs on high-error ONT and
large amplicon inputs (with O(n) placement). - Assorted CLI fixes:
--order anyno longer aborts on a truncated FASTQ that
preserverejects,--estimatehandles empty input,--interleaved 0is
rejected, and--verifyhonors--threads. - Python binding tests now build their own tiny
.fqxvfixture with the CLI
instead of latching onto whatever archive sits at the repo root (the suite ran
for minutes against a large sample and could fail on a half-written one); they
now finish in well under a second. TheInforepr also shows the container
version asformat=major.minorinstead of the raw packed integer.
Security
- Decode-path hardening — a fuzzing-driven pass (cargo-fuzz targets over the
public decode entries, plus a weekly fuzz schedule) closed multiple
decompression bombs across the rANS, quality, sequence, name, reorder, and
container decoders: untrusted length/size headers can no longer trigger
unbounded allocation.
Documentation
- Long-read benchmark tables, a format-comparison page, a Python API reference,
and--quality-bin ont|hifidocumentation.