[0.7.1] - 2026-08-25
Fixed
-
Protenix-v2, OpenDDE and OpenDDE-abag were padding the atom axis and not the token axis. At any
token count that is not a multiple of 64, the ragged tail reached the triangle attention as real
key columns rather than masked ones, and both the stock and the fused attention read them: on a
probe, relative error against the aligned answer is 0.914 ragged against 0.038 padded. A
98-residue fold presented 1208 ragged calls out of 1208 on Protenix-v2 and 1216 on OpenDDE, which
carries a third axis in its structural-token refiner at roughly twice the residue count and so is
essentially never aligned. All three axes now pad to a multiple of 64, mask, and slice back;
TT_BIO_PROTENIX_TOKEN_BUCKET=0restores the old path for an A/B and
TT_BIO_PROTENIX_TOKEN_PAD_MULTIPLEoverrides the multiple.Where the pad is 0 the fold is byte-identical, so a token count that is already a multiple of 64
costs nothing and changes no number: at 512 residues both arms return CIF5e404779d791fa8fon
Protenix-v2 and agree to -0.101 s over eight interleaved pairs on OpenDDE. Where the pad is not 0
you pay for the columns you added, and at real fold lengths that is a win rather than a cost:
298 residues round to 320 and run 4.8 % faster on Protenix-v2 and 6.0 % faster on OpenDDE. Short
folds are the exception and they are worse, because the fold is dispatch-bound before the pad
arrives: 20 residues round to 64 and lose about 17 % of throughput. That is the smoke input
perf_regression.pytimes, so its baseline rows were re-recorded deliberately rather than
waived.docs/size-generality.mdhas the size-by-size reading. -
A job that never got the card now says so instead of looking like a bad result. A co-tenant
holding a device used to surface asRuntimeError: every local worker exited, and inside a gate
as an accuracy verdict against a fold that never ran. Contention now has its own reserved exit
code, 75, carried from the worker through the predict fan-out to the CLI. -
opendde-trpcage-nomsareports a real verdict infull_parity_gate.pyagain instead of
BLOCKED-REGEN. The harvest step wrote the regenerated fixture under the source id rather than the
destination id, so the leg never found what it had just produced. -
The benchmarks page. Rows that are hidden no longer reach the charts and tables, a stray script
fragment no longer prints as page text, and design rows state the batch size they were measured
at. A row is only published once every processor column in it has been measured: partial rows
wait inperf/page_rows_pending.jsonandscripts/site_publish_guard.pyenforces it. -
An explicit
$RF3_CKPTis no longer outranked by a stale checkpoint in the legacy cache
directory. The RF3 accuracy cell fell back to~/.cache/tt-bio/rf3/whenever the path it was
given did not exist, so on a host that still had an old copy there it silently folded that copy
instead of failing on the path the caller named.
Added
-
RFD3 block-sparse atom attention, opt-in with
RFD3_BLOCK_SPARSE=1. The atom site's
128-neighbour index is block-sparse rather than row-sparse, so a block of neighbouring query rows
shares a key window narrow enough to run as a batched dense matmul. Worth 2.982 s/design on a
target it fits (6051 atoms, median of 7 rounds, range 2.218-3.988). It stays off by default
because it is target-specific, not because it is unproven: the query block has to divide the
tile-padded atom axis, which at the default block size holds for 288 of the 11745 atom counts
between 256 and 12000. Every other target takes the dense fallback, which is always available and
caps the downside at the dense chain. Not bit-exact against the dense path, and the only
difference is the softmax reduction order. -
RF3 has an accuracy floor at 997 aa, not just at 117.
release_gate.py --model rf3-1024aafolds
7EIP on the device and gates CA-RMSD against the deposited crystal at 4.0 A;full_parity_gate.py
scores the same leg. Measured 1.9687 A, against a reference of 2.0092 A.
Gates
Cut from a tree that passed, on one Blackhole p300c host (tt-quietbox2, Python 3.12.3, ttnn 0.68.0):
- Packaging: all 61 expected data files and all 43 declared runtime dependencies ship in both the
wheel and the sdist and land on disk after a clean install. - Parity (
full_parity_gate.py, 41 legs, 201 min): 34 PASS, 4 GAP that reproduce their committed
values, 1 PASS-caveated, 1 fixture awaiting regeneration, 1 FAIL. No leg drifted, errored or was
skipped. The FAIL isaf2ig-trunk-device, whose floor was recorded on a 13x10 board with the
template on host and is read here on an 11x10 board with it on card; AF2-IG has no CLI path in
this release. - Accuracy floors (
release_gate.py): every arm passed, including Boltz-2, ESMFold2, Protenix-v2,
OpenDDE, OpenFold3, RF3, OpenBind, BoltzGen, RFD3, Nesso-1, ESMC and the OpenDDE antibody-antigen
DockQ floor. PXDesign fit-RMSD 4.909 A against a 15.0 A floor. - Throughput (
perf_regression.py, +-15%): 14 of 14 models within tolerance on their recorded
p300c baselines. The three models whose numerics this release changed clear their re-recorded
post-fix baselines: Protenix-v2 +13.9%, OpenDDE -0.3%, OpenDDE-abag +0.6%. - CLI (
ux_regression.py): every surface cleared progress, parse and results/manifest shape.
Not covered, and worth naming rather than dropping: the size ladder has no baseline for this card
type, so it did not run; RF3, Nesso-1 and single-sequence ESMC-300M have no throughput baseline on
this card type; RFD3 block-sparse is dark on the gate's own fixture, so its coverage is its own
tests rather than a gate leg; and tt-bio design --model pxdesign is gated through the library
rather than end to end through the CLI.