v0.13.3 — benchmark truth + taint closure
Repair release from adversarial review — the eval suite's instrument audited itself.
Benchmark truth
- Job completion redefined honestly: output contract + brand_check + compliance + non-vacuous rules. Corrected second-agent result: 20% (1/5), down from the previously published 5/5 (which counted empty-checker acceptance as completion). Correction published in eval/RESULTS.md with the original claim annotated, not erased.
- Benchmark fixture gains a governed-voice overlay (never-say, anchors, tone) so text tasks measure brand transfer, not checker emptiness.
Taint closure
- Hostile extracted color names sanitized at every agent-facing compile sink (runtime unknown-role keys, design-synthesis signals, brand_write briefs); raw names stay quarantined in evidence. Adversarial test proves the runtime-key path.
Auditability
- Committed machine-readable receipts (eval/receipts/) for every published LLM run: commit, package, model, per-task contract status, rule coverage, verdicts, tokens.
Full changelog: https://github.com/Brandcode-Studio/brandsystem-mcp/blob/main/CHANGELOG.md