Skip to content

v0.12.16

Latest

Choose a tag to compare

@sauravpanda sauravpanda released this 28 May 19:33
608f75b

Summary

v0.12.16 reverts the v0.12.15 tool-schema slim. The schema slim was intended to lower cached-prefix cost by hiding alias declarations from the default LLM schema, but the WebBench_READ_v5 eval showed it was net-negative: cost/M tokens nearly doubled while accuracy stayed flat.

Why the revert

v0.12.15 vs v0.12.14 per-task cost telemetry on WebBench_READ_v5 (~88 of 198 tasks sampled per run; the 0.12.8-14 lineage tracked tightly in this band):

Metric v0.12.14 v0.12.15 Δ
Success 137/198 (69%) 134/198 (68%) flat (noise)
Cost/M tokens $0.234 $0.453 +94%
Cache-hit rate 62.7% 19.7% -43pp
Cost/task $0.0557 $0.0769 +38%

Mechanism: shrinking the cached prefix invalidated Gemini's implicit cross-task cache. Cache-rebuild cost exceeded the per-call schema savings. Same compensation pattern surfaced in the v0.11.x cycle — any change inside the cached prefix forces a cross-task cache rebuild whose cost exceeds the byte savings.

What's in this release

  • Pure git revert of merge 70a10cb (PR #22).
  • Version bumped 0.12.14 → 0.12.16 across Cargo.toml, Cargo.lock, pyproject.toml.
  • docs/v0.12-plan.md updated:
    • v0.12.15 attempt + result + REJECTED reason.
    • v0.12.16 = pure revert (no code lever; substrate sanity-check).
    • Next slices pivoted away from cached-prefix trims toward per-step uncached bytes.

Validation

  • PYTHONPATH=python pytest tests/test_agent_compat.py tests/test_final_answer_guards.py tests/test_blocked_search_guards.py -q — all 99 tests pass.

Expected eval signature

bu-rust 0.12.16 + WebBench_READ_v5 should reproduce v0.12.14's profile (~69% accuracy, ~$0.234/M tokens, ~63% cache hit). A match confirms the v0.12.15 mechanism and clears the substrate for the next experiment.