Skip to content

v0.11.0 — Trust & Cost: Token Budget, Quality Score, ROI Nudge, Show, Harm Prototype

Choose a tag to compare

@codenamev codenamev released this 30 Apr 21:37
· 151 commits to main since this release
Immutable release. Only release title and notes can be modified.

Theme: Trust & Cost — five user-visible signals that answer "is memory still worth it?" with numbers a skeptical user can read in <30 seconds.

Added

  • Token budget telemetry — every successful SessionStart context injection now records an estimated context_tokens count on its activity_events row. Surfaced three ways:
    • Dashboard Trust panel emits a token_budget block with p50/p95/avg/sample_size over the last 30 days, so the JSON dashboard endpoint and any downstream consumer answer "what does memory cost per session?"
    • claude-memory digest includes a "Context cost" subsection between activity and new-knowledge so the weekly report shows the price tag next to the value.
    • claude-memory stats --tokens [--since DAYS] reports total sessions, p50/p95/avg/min/max, and a histogram across <500 / 500-1k / 1-2k / 2-5k / 5k+ buckets.
  • Pure additive — no schema migration. Historical events written before this release simply contribute zero samples until new injections accumulate.
  • First 0.11.0 milestone item from the 1.0 punchlist (Trust & Cost). Closes the "what % of my SessionStart token budget does memory consume?" gap.
  • Hallucination rate metric — the dashboard now quantifies how clean the fact base is, not just how full it is. Distill::BareConclusionDetector is the production-side mirror of the SessionStart prompt's reason-clause requirement (decision/convention facts must embed "because…" / "so that…" / "to avoid…"). Surfaced two ways:
    • Dashboard Trust panel emits a quality_score block aggregating across project + global active facts: suspect_count (predicate=reference, retagged by ReferenceMaterialDetector), bare_conclusion_count, percentages, and an overall 0–100 score (higher = cleaner). Returns 100 on empty stores so fresh installs aren't penalized.
    • claude-memory digest includes a "Quality" section showing the score breakdown plus the in-window rejection rate ("of facts created in the last 7 days, X% have been rejected since"), so calibration drift is visible.
  • Second 0.11.0 milestone item. Pairs with token-budget telemetry to answer "is memory still worth its cost?" via two skeptic-friendly numbers.
  • claude-memory show — new CLI command prints what memory would inject at the next SessionStart in plain Markdown. Runs the exact Hook::ContextInjector path real sessions use, so output matches what Claude actually receives. Footer reports fact count, ~token estimate, and char count so users see the SessionStart cost at a glance.
    • Default suppresses the raw-transcript "Pending Knowledge Extraction" dump (intended for LLM distillation, not human reading); pass --pending to include it.
    • --source SOURCE (startup/resume/clear) simulates each fresh-session entrypoint so users can preview which sections would appear.
  • Third 0.11.0 milestone item. Closes the inspectability gap — trust requires being able to see what memory will inject, the same way cat CLAUDE.md works.
  • First-week ROI nudge — at SessionEnd, memory now prints memory contributed N facts this session, %used = X for the first 10 sessions, then quiets. New users get user-visible proof memory is doing work for them without having to know about the dashboard. Once trust is established (or it isn't), the nudge gets out of the way.
    • New claude-memory hook nudge subcommand + Hook::Handler#nudge. SessionEnd config now wires [ingest, sweep, nudge] in order.
    • Silent on CLAUDE_MEMORY_NO_NUDGE=1 opt-out, missing session_id, n=0 contributions, and after MAX_NUDGES emissions. The empty-session silent path doesn't burn a slot — quiet sessions don't count toward the 10.
    • Activity event roi_nudge records {n, used, pct, prior_count} per emission so a future migration could change the threshold without re-counting from raw events.
  • Fourth 0.11.0 milestone item. Cold-start trust signal that pairs with #47 (token cost) and #48 (quality) to make the first-week answer to "is this worth it?" visible without effort.
  • Harm benchmark prototype — spec/benchmarks/dataset/harm_scenarios.yml + spec/benchmarks/e2e/harm_bench_spec.rb. Three hand-written cases spanning the riskiest harm classes (stale_tech, mismatched_scope, superseded_undetected). The first ClaudeMemory benchmark that measures whether memory can make Claude wrong — every other benchmark only measures whether memory helps.
    • Structure validation (regex compile, fact loadability, harm-class coverage) runs in stub mode as part of :benchmark tag.
    • Real-mode runner: EVAL_MODE=real bundle exec rspec spec/benchmarks/e2e/harm_bench_spec.rb — needs claude CLI on PATH, ~$2-8 per run. Reports harm rate; doesn't enforce a threshold yet (that's the 0.12 release gate).
  • 0.11.0 risk-de-risking item. If even one of these three surfaces a harm now, the full 10-15-case benchmark planned for 0.12 will likely reveal a fundamental issue — better to learn that at 0.11 than at 0.12. Real-mode prototype run on 2026-04-30 reported 0/3 harm — green light to expand to the full corpus in 0.12.

Changed

  • Hallucination-rate metric calibration — Dashboard::Trust#quality_score now reports a windowed (last 30d) "live" score as the headline plus a "historical" block over all active facts. Production verification on 2026-04-30 (recorded in docs/quality_review.md) showed the unwindowed metric was technically correct but pragmatically misleading: 97% of bare-conclusion facts pre-dated the 2026-04-20 reason-clause prompt commit, and the entire 7-day rejection cluster was a single-class systemic failure (a /study-repo burst), not ongoing noise. The split makes the metric actionable: live score = ongoing extraction quality, historical = legacy data. The digest's "Quality" section uses the live score as the headline.

Fixed

  • Real-eval CLI runner now passes allowed_tools through explicitly so the harm benchmark and other real-mode benches can pre-allow MCP memory tools without per-test wiring.

Upgrade Notes

  • No schema migration. All new features ship purely additive.
  • Hooks run the installed gem from PATH, not the working tree. After upgrading, bundle exec rake install (or gem install claude_memory) is required for the new SessionEnd nudge, claude-memory show command, --tokens stats flag, and context_tokens activity-event field to actually fire on real hook events.
  • Existing quality_score consumers will see additional fields (window_days, historical) in the snapshot. The original keys (score, total_active, suspect_count, bare_conclusion_count, suspect_pct, bare_pct) remain at the top level and now reflect the 30-day live window — historical numbers move to the historical sub-hash.

🧪 Real Eval Validation

Results: 4/6 passed ⚠️ 2 failed
Duration: 73.33s
Estimated Cost: ~$0.12

⚠️ Some real eval tests failed. Check the workflow logs for details.