Releases: antigenomics/arda
Release list
2.36.0 — a junction naming neither V nor J gets a locus proposed too
2.34.0 proposed a missing V or J from the junction, but only within the locus the other side named. A record naming neither came back refused — no locus, no boundary, nothing for a nucleotide stage to work on.
There is no new rule here. Every locus the organism ships now competes, scored by the same anchor_depth from each germline's own anchor, and a locus only wins by explaining residues at both ends — which is what keeps a TRA junction out of TRD, where the V genes are shared and the J genes are not.
Measured. 522 of VDJdb's curated chunks records name neither side (461 distinct keys). All 461 now get a locus and 459 come back good, against none before. The proposed locus agrees with the cdr3.alpha / cdr3.beta column the record was filed under on 457 of 461 (99.13 %) — 318 TRB, 139 TRA. All four disagreements are CACD…DKLIF: TRDV2's own anchor and TRDJ1's own ending, in a schema with no δ column.
The floor is measured too. Each end must explain its own anchor residue or no locus is named. Without it a junction agreeing with nothing (QQQQQQQQQQQQ) reached depth 0 on both ends of every locus and was still handed TRA. Over the 461 real keys the winning locus clears it on every one — 16 at depth 1, 200 at 4, 156 at 5.
scripts/audit_cdr3fix.py now keeps blank-side keys. It required both a V and a J to be named, so the A/B instrument was blind to exactly the rows 2.34.0 and this release change. Corpus 189,596 → 192,726 keys, digest rebases to 0df8541ed1060fe1; on the 189,596 the old filter kept, the digest is unchanged at 2b75491b380310aa — nothing that already worked moved.
Full changelog: https://github.com/antigenomics/arda/blob/master/CHANGELOG.md
v2.35.0 — arda.hmm is the B-cell entry point, not dead code
arda.hmm is no longer deprecated. 2.33.0 deprecated it on the grounds that nothing consumes it
and vdjtools.model.infer_nt_batch answers the same question faster. That reasoning surveyed the
T-cell path and missed the shm= parameter — so it recommended a model with no
somatic-hypermutation term as the replacement for the only SHM-aware one.
Without an SHM model the templated V length is bounded by an exact common prefix, so a single
substitution in the V tail forces the whole rest of it to be re-read as N region.
tests/unit/test_shm_lattice.py has pinned this on IGHV3-30*18 / IGHJ4*02 all along: del_v == 0
with a model against del_v >= len(v_nt) - 3 without one. For a hypermutated IGH junction that is
the normal case, not a corner case, and the recommended replacement prices a mutated V tail as
insertion in exactly the same way the exact-prefix bound does.
The deprecation warning is gone, the roadmap retirement item is gone, and the module docstring now
leads with what it is for. d_prior.tsv stays with it — both of its consumers are in arda
(scenarios._Model and arda.hmm.model_for), and one of them is that scorer.
Unchanged, and still recorded: arda.hmm is not on the annotation path, and the two measured
negatives behind that stand — re-ranking nucleotide D candidates by a scenario likelihood changes
nothing, and replacing the E-value gate with a Bayes factor would need a per-locus shipped threshold.
Both are statements about germline TCR junctions, where an alignment already settles it.
Also — references repaired. vdjtools 4.8.0 ships one D estimator, so
vdjtools.model.posterior_d / posterior_d_batch / load_d_prior no longer exist. Eight places
here still named them; a reader following any of them landed on an ImportError. They now point at
vdjtools.model.annotate_junctions for the amino-acid question and at arda.hmm for the nucleotide
one.
Full changelog: https://github.com/antigenomics/arda/blob/master/CHANGELOG.md
v2.34.1 — walk a germline run once
A packaging release: identical answers, less time, and two invariants that used to hold only by
construction are now tests.
_extend walks a germline run once instead of three times. It is the hottest function in
markup_batch — 159,240 calls per 20,000 junctions — and it used to walk the same span to find the
run, then again to count agreement, then a third time to measure the leading exact prefix. The
offsets it agrees at are already collected on the first walk.
16,287 → 18,338 junction keys/s, so VDJdb's whole 189,596-key corpus marks up in 10.3 s in one
process. The audit decision digest over every one of those keys is unchanged at 2b75491b380310aa,
which is the check that matters: a count-equal, decision-different leg is invisible in a summary.
What a boundary promises is now pinned. Every residue v_end / j_start credits is germline of
the allele they name — measured at 185,636 of 185,636 V keys and 187,272 of 187,272 J keys that place
one, zero disagreements — and cdr3_repaired is a fixed point of itself on all 324 keys that take a
FixAdd. That is the guarantee a consumer applying the repaired junction relies on (#140).
And a functional segment's anchor. A functional V opens with Cys104 and a functional J closes with
Phe or Trp118 on 3,129 of the 3,130 functional entries carrying a templated region across all five
shipped organisms. The single exception is real germline: TRAJ35*01 is IMGT-functional and its
anchor codon is TGC. The test names it rather than repairing it away — asserting there are no
exceptions, or filtering on a [FW]GXG motif, deletes a functional gene from the vocabulary and every
read from it is then called as the nearest J that survived (#138).
Full changelog: https://github.com/antigenomics/arda/blob/master/CHANGELOG.md
v2.34.0 — a blank V or J call is proposed from the junction
A submission is allowed to leave one side out, and a blank is not a reason to refuse the record.
The locus comes from the side that is named, so the proposal is always within the right locus, and
the junction is the evidence about the missing one.
Over VDJdb's 192,726 distinct curation keys, 3,130 leave a side blank (644 with no V, 2,947 with no
J). 2,532 of those now get a segment and 2,504 come back good with both boundaries placed. The 598
that stay refused name neither side, so there is no locus to propose within.
Cdr3Markup.proposed (and the proposed column) says which side was never curated — a different
fact from the allele flag, which means the submission named a different allele of the same gene.
An unresolvable call is still refused: TRBVnope*01 keeps its FailedBadSegment. A submission
that names something wrong has a defect a curator must see; one that names nothing has a gap the
junction can fill.
This closes the last reason a consumer had to carry its own segment proposer alongside cdr3fix.
Full changelog: https://github.com/antigenomics/arda/blob/master/CHANGELOG.md
2.33.0 — the repair policy, and the D posterior moves to vdjtools
Closes #141, #142, #143, #144.
The repair policy: when it is not certain, flag it — do not change it
cdr3fix admits four edits, in this order, and rewrites nothing else:
- re-call the V or J when another allele explains 2 more residues contiguously from its own anchor and does not break an anchor the submission already holds;
- trim framework past that allele's anchor;
- add back germline the submission was cut inside of;
- substitute the first or last residue only, to the anchor, on a solid match.
A residue disagreeing with germline inside the templated run is reported (new mismatch flag) and left as submitted — a curation error and an allele IMGT does not record are indistinguishable from one junction. The previous behaviour turned CAAAETSYDKV{M,R,T,V}F on TRAJ50 into four copies of CAAAETSYDKVIF.
_RECALL_GAIN = 2 is measured against isalgo/airr_control's nucleotide-called V/J on 3,059 human TRA and 426 human TRB curation keys: the smallest margin at which no re-call moves a call away from the nucleotide answer.
Against the authoritative VDJdb 2026-06-03 release, all 187,488 keys
| 2.32.0 (unreleased) | 2.33.0 | |
|---|---|---|
| repaired junction agrees | 98.3604 % | 99.8352 % |
| the release's own repairs reproduced | 4,267/4,499 | 4,331/4,499 |
| rewrites the release never ships | 2,842 | 141 |
| — into an already-canonical junction | 2,744 | 35 |
| non-canonical junctions shipped (release: 715) | 709 | 677 |
arda ships fewer malformed junctions than the release it is measured against. Reproduce with scripts/compare_vdjdb_release.py.
BREAKING: arda.dpost moved to vdjtools.model.posterior_d
A posterior over the D gene marginalises model distributions, so it belongs with the recombination model (#144). Ported, not rewritten: identical field-for-field on 3,000 real human TRB junctions, and it gained the batch entry point #142 asked for. arda markup --d-posterior / --d-prior are gone. The prior table, its fitter and its new reader (arda.scenarios.load_prior_table) stay here.
Needs vdjtools>=4.8.0 for the posterior.
Also
v_alts/j_alts— the alleles a junction cannot separate travel with the answer, for the nucleotide stage to settle.map_d_junction(v_end=, j_start=)— take the interior instead of re-deriving it from an inferred sequence.arda.hmmdeprecated: nothing consumes it, andvdjtools.model.infer_nt_batchanswers the same question batched and in C++.- #143 is closed by the 2.32.0 port — there is no Python alignment DP in
cdr3fixany more; throughput is 28,900 keys/s including the new per-record re-call.
1,312 tests pass.
arda-mapper 2.31.0 — the germline boundary in nucleotides
Cdr3Markup gains v_end_nt and j_start_nt beside v_end and j_start. The residue counts are
unchanged — dpost slices the non-templated middle with them, and that contract is what they are for.
An exonuclease does not stop on a codon boundary, so the last residue a germline touches is usually
part germline and part N region, and an alignment on the protein can only round that to a whole
residue. What the protein still fixes is which nucleotides are admissible: 17 of the 20 residues open
every one of their codons with the same base. Against a germline GGA (Gly) an observed Glu can only
be GAA or GAG, both of which begin with the germline's G, so the germline demonstrably reaches
one nucleotide further — and an aligner reading the observed sequence counts it as germline whether it
was templated or an insertion reproduced it. Where a residue's codons disagree, they are weighed by
how far each would let the germline reach: a templated nucleotide is free, one that matches by chance
costs 1/4.
Measured against isalgo/airr_control's human.trb.ntvj, on the 8,133 VDJdb human TRB junctions
whose boundary every control observation agrees on:
| quantity | 2.30.1 | 2.31.0 |
|---|---|---|
v.end, VDJdb residue convention (nt + 1) // 3, exact |
5,839 (71.79 %) | 7,556 (92.91 %) |
j.start, VDJdb residue convention ceil(nt / 3), exact |
7,968 (97.97 %) | 7,968 (97.97 %) |
| V boundary in nucleotides, exact | 2,355 (28.96 %) | 6,538 (80.39 %) |
| J boundary in nucleotides, exact | 2,462 (30.27 %) | 6,057 (74.47 %) |
VDJdb's legacy k-mer scanner measures 71.8 % on the same set — the same number as 2.30.1, because both
answer in whole residues. j.start cannot move and does not: VDJdb defines it as the first fully
J-templated residue, which is the same residue whether the germline reaches one or two nucleotides into
the one before it.
The two sides are bounded by different things. The V residue count is right on 83.99 % of records and
the extension on 95.71 % of those, so V is limited by the protein alignment. The J count is right
on 97.97 % and the extension on 76.02 %, because the nucleotide a J boundary turns on is the
codon's third — the one position the genetic code leaves free.
This is antigenomics/vdjtools#182; vdjtools 4.7.0 consumes it as model.germline_boundary.
arda 2.30.1 — the amino-acid V/J boundary stops where germline evidence stops
cdr3fix placed the V/J boundary about two residues too far into the junction on amino-acid
input. _align picks the best-scoring cell and, on a tie, preferred consuming more query - but
a tie means those residues contributed nothing, so crediting them to the segment is attribution
without evidence. The scoring is nucleotide calibration (_MATCH, _MISMATCH = 1, -1), where one
substitution is a plausible read error; on amino acids one coincidental match pays for a mismatch
exactly and the alignment walks into the N region.
The fix separates two questions that shared one answer. _errors and _repair keep the full
aligned extent, so a substitution one residue short of the end is still reported and still
repaired. v_end and j_start come from _supported, the first position where the running
score reaches its maximum; past that point the alignment is only breaking even. A repair earns the
full extent, because a repaired residue is one the germline explains.
Measured against observed nucleotide markup of real repertoires (boundaries established with the
nucleotides present; arda seeing only amino acids and the V/J calls):
v.end exact |
within 1 | j.start exact |
within 1 | out by >= 2 | |
|---|---|---|---|---|---|
| 2.29.0, 25,000 human TRB | 72.6 % | 98.7 % | 95.9 % | 97.7 % | 566 on j.start |
| 2.30.1 | 73.3 % | 99.3 % | 97.7 % | 99.5 % | 10 on j.start |
On the 8,334 VDJdb human TRB records that also appear in that data, j.start exact goes
8,041 -> 8,164, overtaking the k-mer scanner VDJdb currently ships (7,987) while declining
nothing that scanner declines. Over-extension by two residues or more: 138 -> 1.
1,213 tests pass and no test was modified, including the deep-substitution case the two-extent
split exists to preserve.
Closes #132. What the fix does not do is recorded in #134 and ROADMAP.md item 12: a
substitution-matrix score for the 170 records still off by one residue, the TRBD2/TRBJ1 locus
constraint the D caller ignores, and the 658 sides where a sibling allele of the same gene explains
the junction with no edit at all.
⚠ 2.30.0 was tagged but never published: its tree fails its own version-pin tests. 2.30.1 is that
release with the version carried into all six places that hold it.
arda 2.29.0 — a model of where hypermutation is expected
Added: arda shm-model — where a mutation is EXPECTED
arda.shm reports where a read differs from its germline. This fits the other half:
P(substitution | 5-mer germline context, region), estimated from what every mode already writes
under the default --shm framework.
arda shm-model -i mapped.airr.tsv -o shm.tsv --locus IGH⛔ Not a per-allele-per-position table, and the measurement is why. Across two donors a
per-allele-per-position profile transfers at Pearson r .2557–.5601 against 5-mer context's
r .7424–.7854; within one donor the positional table repeats at r ≈ .99, because it is a
portrait of that donor's expanded clones rather than a property of the allele. ⚠ Context is also
the only one of the two that reaches the V tail inside the junction — exactly the positions
arda.shm scopes out as unidentifiable, where a position has no estimable rate at all and a
5-mer takes its rate from every allele carrying it in clean framework sequence.
✅ Both parameter families are earned. AID's WRCY/RGYW motifs carry 4.66× / 5.00× / 4.77× /
4.69× the rate of every other covered position across four libraries, on an overall rate that
itself moves 1.6× between donors; and the CDR:FWR residual after conditioning on context
reproduces across donors — FWR1 0.630 / 0.626, CDR2 1.389 / 1.291.
⚠ The fitted scale is the sample's, not the model's, so it is written as provenance and never
applied.
Added: arda scenarios --shm-model — the consumer
scenarios.lattice bounded the templated V length with an exact common prefix, so on IGH one
substitution in the V tail forced the rest of that tail to be explained as insertion — and delV
and insVD are the two distributions the estimator fits. Those positions are now priced:
Π(1 − μ) over matches, μ / 3 over mismatches.
Measured on 26,619 real IGH junctions, with the model fitted on a different donor's library:
| parameter | exact | with --shm-model |
change |
|---|---|---|---|
delV mean |
4.148 | 2.708 | −1.440 |
delV P(no deletion) |
.1426 | .2388 | +.0962 |
insVD mean |
12.065 | 10.972 | −1.093 |
insDJ mean |
12.342 | 12.144 | −0.198 |
delJ mean |
12.247 | 12.253 | +0.006 |
The exact bound was charging 1.44 nt of V germline per IGH rearrangement to deletion, and
delV's mode moves from 1 to 0. ✅ The model reaches the V side only, so delJ at +0.006 against
delV at −1.440 is the result's own control. Cost: +4.8 % wall on a 3-iteration fit.
✅ The exact-match bound turns out to be the μ = 0 case of the new emission, not a separate
rule. shm=None remains the default and is byte-identical, so TR and unmutated IG are untouched.
Fixed: a V/J call naming a FAMILY with one functional gene now resolves
resolve_allele climbed exact → gene*01 → first allele of the gene, and none of those reaches a
call naming a family whose genes all carry a suffix: nothing is named TRBV20, only
TRBV20-1, so markup_cdr3 returned FailedBadSegment and dropped the V end entirely.
6,291 human beta chains in VDJdb carry such a call (TRBV20 1,839, TRBV3 1,039, TRBV24
291, TRBV29 70), and the germline places their CDR3s without a single edit once it resolves.
Never: the new rung refuses rather than guesses — the family must hold exactly one
functional gene. TRBV3 resolves because TRBV3-2 is a pseudogene; TRBV6, with five
functional genes, still returns "".
Fixed: a tie-group v_call silently found no germline
IGHV3-23*01 lives inside IGHV3-23*01,IGHV3-23D*01, so a lookup by a single allele found
nothing — it cost the SHM estimator reads and made the first SHM-aware lattice score exactly
like the exact one, which reads as a clean null rather than a missing key. 325 human IGH keys
become 349.
Full detail in CHANGELOG.md.
arda 2.28.0 — the integrations speak airrflow
Breaking: all three runners speak nf-core/airrflow's vocabulary, not an arda-only one
process ARDA → ARDA_ASSIGN, a drop-in for CHANGEO_ASSIGNGENES + CHANGEO_MAKEDB, and
params.regime is gone — the mode now comes from --library_generation_method. A second
arda-only name for the library protocol was a second place to get it wrong, and getting it wrong
is a silent 2–4× slowdown rather than an error. arda.samples.read_sheet reads both dialects, so
one samplesheet drives Nextflow, Snakemake and arda cluster with no translation step.
Migration table in integrations/nextflow/arda/README.md.
⛔ sc_10x_genomics is refused with a message: arda cells takes one per-molecule UMI consensus
FASTQ with the barcode in the record name, not a raw 10x read pair.
Added
arda markup --d-prior PATH— score against a fitted D prior without overwriting a file
inside the installed database. Implies--d-posterior; a path that is not there raises.arda resolve-ties --loci IGK,IGL— whether widening the V call helps is a property of the
locus. Exactv_gene-set agreement against anarda igblasttruth, eleven arms:
IGK .5707 → .9373, IGL +5.5 pt, TRB +4.0 pt, and IGH −0.6 to −14.9 pt. Named loci are
widened, every other row is copied through untouched.arda igblast --receptor ig|tr|both— a library whose receptor type is known no longer pays
a whole IgBLAST pass over the other one.bothstays the default.
Documented: what an IG V call means when the read is short
v_gene recall against an IgBLAST truth is .1170 on reads covering under 60 nt of V germline and
.9896 on reads covering 200 nt or more — and .9896 is the TRA amplicon's .9867, so there is no
IG-specific deficit at that coverage. Position beats length, somatic hypermutation does not order
this, and no threshold ships. docs/usage.rst.
Fixed
- A
d_priortablearda scenarioswrote could not be read back: the generator emits a#
provenance line above the header andload_d_priorskipped line 1 by position. - The Nextflow Dockerfile's acceptance check grepped
arda rnaseq --helpfor four flags that moved
toarda mapin 2.16.0 — everydocker buildfailed on a correct install. - The Cys104 junction gate no longer discards a junction over one substitution in its first two
bases: +144 correct junctions and one extra over-extension across 92,466 truth junctions,
held-out locus included. build_germline_dbsignoredlocus.v_shared, so IgBLAST never saw a TRDV germline when marking
up chimeric scaffolds: complete markup 7/483 → 49/483 with every V call correct.
Full detail in CHANGELOG.md.
arda 2.27.0 — personalized germline
Added: personalized germline — arda genotype, arda resolve-ties --genotype
A reference catalogues every allele anyone carries; no donor carries all of them and none carries
more than two per gene, so a call naming a third is wrong before any sequence is looked at. TIgGER
measured 11.2 % -> 1.5 % ambiguous V assignments from restricting calls to an inferred genotype on
full-length BCR.
resolve-ties --genotype adds v_call_genotyped and leaves v_call byte-identical. Never: it
re-assigns, it never re-aligns and never rebuilds a reference -- scaffold ids are positional so
changing the allele set renumbers every scaffold in the locus, build-db needs IgBLAST and IMGT
network access a pip install does not have, and the mmseqs freshness contract records no allele-set
identity, so a donor-specific FASTA beside the shared one would silently invalidate the index for
every concurrent process. Given the span a read already aligned over, the restriction is a set
intersection. Any source works -- an OGRDB set, a MiXCR-inferred library, a one-column TSV of
allele names -- and an allele the reference cannot call raises.
Never: an allele survives unless its own gene was genotyped and did not name it. Tie lists
routinely span genes, so a gene-blind test empties every read whose tie list merely brushes a
genotyped gene: measured on the 453-read fixture with a one-gene genotype, 79 rows came back empty
and every one was a read of some other gene.
Never: a row that never had a v_call is counted in no_call, not contradicted -- 74 of that
fixture's 453 rows are J-only, and scoring them as contradictions made a correct restriction look
like it had rejected a sixth of the library.
arda genotype infers the set, and the call is a likelihood ratio between diploid genotypes,
not a coverage rule.
The junction is de facto a UMI. V(D)J junctional diversity makes a nucleotide junction
essentially unique to one rearrangement, so grouping reads by (locus, gene, J gene, junction)
groups them by molecule. Two things follow and the inference rests on both. The clonotype, not
the read, is the unit of observation -- one junction is one independent draw from the donor's two
chromosomes however many reads carry it, so read depth is reported beside the clonotype count and
never substituted for it. And disagreement within a junction is error, not allele -- every read
of one rearrangement carries the same V allele by construction, so a read naming another is
hypermutation or sequencing error, which is where the per-read miscall rate is measured from
instead of being a constant somebody picked (5.25e-04 on the amplicon below).
⚠ Never: that error rate is measured over ALL reads, not the germline-exact subset the assignment
uses. Those were selected for carrying no mismatch, so they never disagree and the estimate
collapses silently onto its own floor -- 0 discordant reads of 77,345 on a real library, i.e. no
estimate at all dressed up as a small one.
Every single allele and every pair is scored by its multinomial likelihood under that rate;
homozygous and heterozygous are the same formula at genotype size 1 and 2, and a gene is called
only when the winner beats the runner-up by --min-log10-bf (default 1.0). Ambiguous clonotypes
still constrain the answer: discarding them is not conservative but wrong -- if a donor is
*01/*02 and only *02 is separable at the read length, every unambiguous observation is *02,
the likelihood says homozygous, and the restriction then empties every *01 read, reproducibly.
⚠ The first version of this was TIgGER's frequency rule and it had to be replaced. "The fewest
alleles explaining 7/8 of the calls" has no error model, so it cannot tell overwhelming evidence
from none: on a real TRB amplicon it called TRBV11-2 off 754 of 757 clonotypes that singled
the allele out and TRBV20-1 off 1 of 2,544, reported both as explained = 1.0000, ok, and
returned 43 of 53 genes with not one heterozygous. A ratio of counts is not a confidence.
⚠ Allele-level genotyping is a read-length feature, and most libraries do not have it. Measured
over the 801 committed human V germlines, the shortest 3'-anchored span separating a gene's alleles
is a median of 150 nt for TRBV (175 TRAV, 230 IGHV); only 16 of 44 multi-allele TRBV genes
separate within 100 nt. End to end on SRR5233641 (human TRB amplicon, 151 nt paired, 99,839
mapped reads -> 33,440 clonotypes / 77,345 reads): reads cover a median of 72 nt of V
germline, 56 nt after clipping at the Cys104 anchor, and not one reaches 150. So 33.1 % of
clonotypes can be assigned an allele and 17 of 53 genes are called -- 11 single_allele
(one catalogued allele, carried without inference) and 6 genuinely inferred, at log10 Bayes
factors of 223 (TRBV11-2), 251 (TRBV5-6) and 11.7 (TRBV5-8). The other 36 are refused,
including TRBV10-3 with 1,037 clonotypes none of which separate its alleles. The refusals are the
point; a full-length, 5'RACE or arda cells library has the resolution this one does not.
Fixed: the restriction report counted reorderings as narrowings
TieResolver.candidates() returns its tie list sorted by name while v_call carries the aligner's
order, so restrict emitted TRAV20*01,TRAV20*02 where the input said TRAV20*02,TRAV20*01 --
the same two alleles -- and the report's string test scored every one of those as a narrowing.
Measured on a 100,000-read TRA amplicon: 20,587 reported, 281 real, 20,306 pure reorderings.
Two fixes. Surviving alleles now keep v_call's own order, so a restriction that removes nothing
returns a byte-identical string; and narrowed counts rows that lost an allele, not rows
whose string changed, with a fourth bucket (recalled) so narrowed + unchanged + contradicted +
recalled partitions the assessed rows exactly. Never: a metric that reports something other than
what its name says is the defect class this repo keeps paying for.
Also: infer_genotype now raises on an unrecognised --scope instead of falling through to
full. The two differ only in where each read's germline span is clipped, so a typo would widen
every span into the junction, narrow every tie set, and return calls more confident than the data
supports -- with no error anywhere.
Changed: the benchmark tables are full-pipeline, three-way, and on one denominator
README.md, docs/usage.rst and the Nextflow module's README carried stage-vs-stage numbers from
2.11.1 -- arda's AIRR-emitting stage against MiXCR's non-AIRR-emitting one. Replaced with
same-job, end-to-end-to-a-clonotype-table runs of arda 2.27.0 vs MiXCR 4.7.0 vs TRUST4, three
reps, medians, each tool at its best preset (benchmark repo, round 26).
- TRA amplicon, 100 k reads -- MiXCR is 1.84x faster on wall in its own regime; arda gets
there on 1.41x less CPU and 3.16x less RSS and returns the most clonotypes (19,841 vs
19,697 vs 18,559) over the most reads (43,503 / 42,712 / 37,688). TRUST4 is 5.1x slower than
arda -- the first same-job end-to-end amplicon wall against TRUST4 in this project. - Bulk RNA-seq, 660 k pairs -- arda returns +27.8 % clonotypes and +97.9 % reads assigned
against MiXCR and +14.0 % / +35.8 % against TRUST4, at 1.38x MiXCR's wall and 2.76x less RSS.
⚠ TRUST4 is genuinely 1.61x faster on wall at 4.0x less CPU here, reaching 87.7 % of arda's
clonotypes and 73.6 % of its assigned reads.
⛔ The accuracy arm now prints coverage before rates. A per-tool inner join gives each tool its
own denominator: of 48,033 IgBLAST truth reads at v_score >= 70, arda emits a row for 48,030
and MiXCR for 46,503, so a joined comparison silently drops 1,530 reads MiXCR never answered.
Over all truth reads arda leads v_gene recall .9867 vs .9660 and precision .9996 vs .9977;
on the common subset MiXCR leads recall .9977 vs .9869 while arda still leads precision .9997 vs
.9977. Both denominators ship.
✅ No regression, measured rather than assumed. arda 2.18.0 from PyPI against this release in
one job, legs alternating, 3 reps, same committed reference (database/ is unchanged since
v2.18.0) and same mmseqs: amplicon 13.94 -> 13.66 s, bulk 21.45 -> 21.74 s (1.4 %, inside
the rep spread and accounted for by the per-run QC stage 2.20.0 added), RSS identical. Clonotype
output is byte-identical on both regimes by call digest, not by row count. And the published
2.11.1 accuracy figures reproduce to every digit fifteen releases on -- v_gene recall .9867,
precision .9996, j_gene .9892 / .9953, junction precision among emitted .99919.
Fixed: arda resolve-ties raised on every plain pip install
Its germlines came from the IMGT source tree, which is arda build-db's input and ships with
neither the wheel nor the reference tarball -- the same shape as the segments.fasta deploy trap,
and invisible in a source checkout because one is sitting right there. The new
arda.germline.segment_germlines derives them from the committed markup.tsv + alleles.fasta
instead, which fixes two more things at once: the coordinate spaces now agree (scaffolds are
trimmed to coding frame, so every v_germline_start/_end arda emits was 1-2 nt off against the
raw IMGT sequence for 10 of 884 human V alleles), and tie lists can no longer offer an allele the
reference cannot call (884 functional in IMGT, 801 reach a scaffold).
Never: markup.tsv's v_call is not always one allele. Scaffolds are deduplicated by
assembled sequence, so 23 human V entries are comma-joined groups hiding 49 alleles (3 more for J).
Keyed on the group string they are invisible to every consumer -- TieResolver.expand takes
call.split(",")[0], misses the group key, and returns the call untouched, a tie list that
silently never fires. Split into members: 801 V, 132 J.
Never: v_sequence_end i...