Per-model golden sets from the released artifacts: 12 models x 25 inputs, sha-pinned. The corpus behind every runtime's byte-parity CI.
Parity is precision-scoped (measured 2026-09-07): fp32/fp16 rows are byte-assertions — identical on every runtime and hardware tested. int8/int4 rows are reference outputs with documented provenance: quantized decode flips near-tie decisions across CPU architectures, so those artifacts are quality-gated (per-model cer_delta) rather than byte-gated. Runtimes assert decode health on quantized rows and byte equality on the rest.