Skip to content

v0.4.0 — pairwise alignment without BioPython, and two cache-corruption fixes

Choose a tag to compare

@mikessh mikessh released this 14 Jul 21:18
· 37 commits to master since this release

Pairwise alignment without BioPython — plus two cache-corruption fixes that were tagged 0.3.1 but never published.

Added

  • seqtree.pairwise — Needleman–Wunsch and Smith–Waterman. Ordinary protein alignment on the raw log-odds scale, so reaching for BioPython is no longer necessary.

    pairwise.score(q, r, matrix, mode=...) optimal score, O(min(m,n)) memory
    pairwise.align(...) plus the aligned strings and ops
    pairwise.score_matrix(queries, refs, ...) dense n × K, GIL released, zero-copy numpy
    pairwise.dist_matrix(...) d = s(a,a) + s(b,b) − 2·s(a,b) — non-negative, zero on the diagonal

    mode="global" is Needleman–Wunsch, mode="local" is Smith–Waterman, and gap_open == gap_extend gives linear gaps — no separate mode. A gap run of length L costs gap_open + (L-1)·gap_extend, and global charges end gaps (true NW, not semi-global).

    It is a drop-in. Verified against Bio.Align.PairwiseAligner as an oracle across three matrices × fifteen gap/mode settings × sixty sequence shapes — zero disagreements — and on real germline V genes. BioPython is a test-only dependency; seqtree still has zero required runtime dependencies and never imports it.

    And faster, because there is no Python in the per-pair loop:

    sequence length seqtree, 1 thread seqtree, 16 threads BioPython speedup
    15 (a junction) 1.7 M pairs/s 20.1 M pairs/s 0.31 M pairs/s 65×
    90 (a germline V gene) 72 k pairs/s 893 k pairs/s 10 k pairs/s 87×
  • SubstitutionMatrix.similarity(a, b) — the raw signed log-odds, alongside the existing non-negative penalty(a, b). The Gram transform pen = s(a,a) + s(b,b) − 2·s(a,b) is lossy (it forces the diagonal to zero), so both views are now stored.

  • SubstitutionMatrix.blosum45() and .blosum80(), usable wherever a matrix name is accepted.

Fixed

  • A cold cache shared by concurrent processes could hand back a half-written index. Index::save wrote straight into the destination, so a process that checked for the cache while another was still writing it loaded a truncated file and raised truncated or corrupt index — 10 times out of 10 on the 45 MB control. Saves now write a temporary and rename it into place (atomic on the same filesystem, POSIX and Windows alike).

    This is the first-use failure of any multi-process fan-out sharing ~/.cache: pytest-xdist, a Snakemake or Nextflow rule, a multiprocessing pool. Warm caches were never at risk, and CI matrix jobs (separate runners) were never affected.

  • The control cache is content-addressed, so a stale cache can no longer be served silently. The old key named neither the alphabet, nor the seed, nor the source data — so two calls that must draw different samples shared one file, and an upgrade that changed the bundled control served the previous release's control from a warm cache. You no longer need to clear ~/.cache/seqtree when upgrading.

  • Index.align compared raw characters instead of codec-encoded ones, so a lowercase query against an identical uppercase reference scored 12 under unit cost but 0 under a matrix, and both labelled identical residues as substitutions.

Full detail in CHANGELOG.md.