Repository navigation
Summary
v0.12.16 reverts the v0.12.15 tool-schema slim. The schema slim was intended to lower cached-prefix cost by hiding alias declarations from the default LLM schema, but the WebBench_READ_v5 eval showed it was net-negative: cost/M tokens nearly doubled while accuracy stayed flat.
Why the revert
v0.12.15 vs v0.12.14 per-task cost telemetry on WebBench_READ_v5 (~88 of 198 tasks sampled per run; the 0.12.8-14 lineage tracked tightly in this band):
| Metric | v0.12.14 | v0.12.15 | Δ |
|---|---|---|---|
| Success | 137/198 (69%) | 134/198 (68%) | flat (noise) |
| Cost/M tokens | $0.234 | $0.453 | +94% |
| Cache-hit rate | 62.7% | 19.7% | -43pp |
| Cost/task | $0.0557 | $0.0769 | +38% |
Mechanism: shrinking the cached prefix invalidated Gemini's implicit cross-task cache. Cache-rebuild cost exceeded the per-call schema savings. Same compensation pattern surfaced in the v0.11.x cycle — any change inside the cached prefix forces a cross-task cache rebuild whose cost exceeds the byte savings.
What's in this release
- Pure git revert of merge
70a10cb(PR #22). - Version bumped 0.12.14 → 0.12.16 across
Cargo.toml,Cargo.lock,pyproject.toml. docs/v0.12-plan.mdupdated:- v0.12.15 attempt + result + REJECTED reason.
- v0.12.16 = pure revert (no code lever; substrate sanity-check).
- Next slices pivoted away from cached-prefix trims toward per-step uncached bytes.
Validation
PYTHONPATH=python pytest tests/test_agent_compat.py tests/test_final_answer_guards.py tests/test_blocked_search_guards.py -q— all 99 tests pass.
Expected eval signature
bu-rust 0.12.16 + WebBench_READ_v5 should reproduce v0.12.14's profile (~69% accuracy, ~$0.234/M tokens, ~63% cache hit). A match confirms the v0.12.15 mechanism and clears the substrate for the next experiment.