Repository navigation
v0.11.0 — Trust & Cost: Token Budget, Quality Score, ROI Nudge, Show, Harm Prototype
·
151 commits
to main
since this release
Immutable
release. Only release title and notes can be modified.
Theme: Trust & Cost — five user-visible signals that answer "is memory still worth it?" with numbers a skeptical user can read in <30 seconds.
Added
- Token budget telemetry — every successful SessionStart context injection now records an estimated
context_tokenscount on itsactivity_eventsrow. Surfaced three ways:- Dashboard Trust panel emits a
token_budgetblock with p50/p95/avg/sample_size over the last 30 days, so the JSON dashboard endpoint and any downstream consumer answer "what does memory cost per session?" claude-memory digestincludes a "Context cost" subsection between activity and new-knowledge so the weekly report shows the price tag next to the value.claude-memory stats --tokens [--since DAYS]reports total sessions, p50/p95/avg/min/max, and a histogram across <500 / 500-1k / 1-2k / 2-5k / 5k+ buckets.
- Dashboard Trust panel emits a
- Pure additive — no schema migration. Historical events written before this release simply contribute zero samples until new injections accumulate.
- First 0.11.0 milestone item from the 1.0 punchlist (Trust & Cost). Closes the "what % of my SessionStart token budget does memory consume?" gap.
- Hallucination rate metric — the dashboard now quantifies how clean the fact base is, not just how full it is.
Distill::BareConclusionDetectoris the production-side mirror of the SessionStart prompt's reason-clause requirement (decision/convention facts must embed "because…" / "so that…" / "to avoid…"). Surfaced two ways:- Dashboard Trust panel emits a
quality_scoreblock aggregating across project + global active facts:suspect_count(predicate=reference, retagged by ReferenceMaterialDetector),bare_conclusion_count, percentages, and an overall 0–100 score (higher = cleaner). Returns 100 on empty stores so fresh installs aren't penalized. claude-memory digestincludes a "Quality" section showing the score breakdown plus the in-window rejection rate ("of facts created in the last 7 days, X% have been rejected since"), so calibration drift is visible.
- Dashboard Trust panel emits a
- Second 0.11.0 milestone item. Pairs with token-budget telemetry to answer "is memory still worth its cost?" via two skeptic-friendly numbers.
claude-memory show— new CLI command prints what memory would inject at the next SessionStart in plain Markdown. Runs the exactHook::ContextInjectorpath real sessions use, so output matches what Claude actually receives. Footer reports fact count, ~token estimate, and char count so users see the SessionStart cost at a glance.- Default suppresses the raw-transcript "Pending Knowledge Extraction" dump (intended for LLM distillation, not human reading); pass
--pendingto include it. --source SOURCE(startup/resume/clear) simulates each fresh-session entrypoint so users can preview which sections would appear.
- Default suppresses the raw-transcript "Pending Knowledge Extraction" dump (intended for LLM distillation, not human reading); pass
- Third 0.11.0 milestone item. Closes the inspectability gap — trust requires being able to see what memory will inject, the same way
cat CLAUDE.mdworks. - First-week ROI nudge — at SessionEnd, memory now prints
memory contributed N facts this session, %used = Xfor the first 10 sessions, then quiets. New users get user-visible proof memory is doing work for them without having to know about the dashboard. Once trust is established (or it isn't), the nudge gets out of the way.- New
claude-memory hook nudgesubcommand +Hook::Handler#nudge. SessionEnd config now wires[ingest, sweep, nudge]in order. - Silent on
CLAUDE_MEMORY_NO_NUDGE=1opt-out, missing session_id, n=0 contributions, and after MAX_NUDGES emissions. The empty-session silent path doesn't burn a slot — quiet sessions don't count toward the 10. - Activity event
roi_nudgerecords{n, used, pct, prior_count}per emission so a future migration could change the threshold without re-counting from raw events.
- New
- Fourth 0.11.0 milestone item. Cold-start trust signal that pairs with #47 (token cost) and #48 (quality) to make the first-week answer to "is this worth it?" visible without effort.
- Harm benchmark prototype —
spec/benchmarks/dataset/harm_scenarios.yml+spec/benchmarks/e2e/harm_bench_spec.rb. Three hand-written cases spanning the riskiest harm classes (stale_tech, mismatched_scope, superseded_undetected). The first ClaudeMemory benchmark that measures whether memory can make Claude wrong — every other benchmark only measures whether memory helps.- Structure validation (regex compile, fact loadability, harm-class coverage) runs in stub mode as part of
:benchmarktag. - Real-mode runner:
EVAL_MODE=real bundle exec rspec spec/benchmarks/e2e/harm_bench_spec.rb— needsclaudeCLI on PATH, ~$2-8 per run. Reports harm rate; doesn't enforce a threshold yet (that's the 0.12 release gate).
- Structure validation (regex compile, fact loadability, harm-class coverage) runs in stub mode as part of
- 0.11.0 risk-de-risking item. If even one of these three surfaces a harm now, the full 10-15-case benchmark planned for 0.12 will likely reveal a fundamental issue — better to learn that at 0.11 than at 0.12. Real-mode prototype run on 2026-04-30 reported 0/3 harm — green light to expand to the full corpus in 0.12.
Changed
- Hallucination-rate metric calibration —
Dashboard::Trust#quality_scorenow reports a windowed (last 30d) "live" score as the headline plus a "historical" block over all active facts. Production verification on 2026-04-30 (recorded indocs/quality_review.md) showed the unwindowed metric was technically correct but pragmatically misleading: 97% of bare-conclusion facts pre-dated the 2026-04-20 reason-clause prompt commit, and the entire 7-day rejection cluster was a single-class systemic failure (a/study-repoburst), not ongoing noise. The split makes the metric actionable: live score = ongoing extraction quality, historical = legacy data. The digest's "Quality" section uses the live score as the headline.
Fixed
- Real-eval CLI runner now passes
allowed_toolsthrough explicitly so the harm benchmark and other real-mode benches can pre-allow MCP memory tools without per-test wiring.
Upgrade Notes
- No schema migration. All new features ship purely additive.
- Hooks run the installed gem from PATH, not the working tree. After upgrading,
bundle exec rake install(orgem install claude_memory) is required for the new SessionEnd nudge,claude-memory showcommand,--tokensstats flag, andcontext_tokensactivity-event field to actually fire on real hook events. - Existing
quality_scoreconsumers will see additional fields (window_days,historical) in the snapshot. The original keys (score,total_active,suspect_count,bare_conclusion_count,suspect_pct,bare_pct) remain at the top level and now reflect the 30-day live window — historical numbers move to thehistoricalsub-hash.
🧪 Real Eval Validation
Results: 4/6 passed
Duration: 73.33s
Estimated Cost: ~$0.12