Skip to content

v5.12.0 — The Proof Edition

Choose a tag to compare

@Wolfe-Jam Wolfe-Jam released this 22 Jun 04:10

The Proof Edition

faf_bench proves FAF's grounding lift in-session — a cold-vs-.faf benchmark with a mechanical receipt — and faf_go now bootstraps from a cold repo (init → auto → 6Ws).

Added

  • faf_bench — the in-session AI-grounding benchmark, promoted to lead the Core tier. Derives a question set from the .faf, has the model answer cold (no context) and with the .faf, grades mechanically (the .faf is the answer key), and emits a receipt showing the delta. Proof you can run, in-session — not a pitch.
  • /faf-bench prompt — the honest two-pass session protocol (cold-blind first; never invents token counts).

Changed

  • faf_go bootstraps a cold repo — no project.faf? it runs init → auto first, then the 6Ws.
  • Core tier is now 13 (faf_bench leads; faf stays Extended to keep the Glama-scored default surface coherent — AAA protected).

36 tools (13 Core, rest via FAF_TOOLS=all). 569 tests, 0 fail.

FAF defines. MD instructs. AI codes. 🏎️