Skip to content

v0.12.0 — Release Discipline, Observability, Self-Audit

Choose a tag to compare

@codenamev codenamev released this 01 Jun 13:17
· 117 commits to main since this release
Immutable release. Only release title and notes can be modified.

Theme: Release Discipline + Observability + Self-Audit — the infrastructure that makes a 1.0 semver promise defensible. This release locks down the public API surface, adds the observability primitives (OTel ingestion, dashboard Telemetry) and the self-audit toolkit (claude-memory audit) that serve the visibility pillar, and ships the negative-fact harm benchmark + staleness guard that make the long-horizon-quality claim measurable rather than aspirational.

Added

  • Staleness guard for single-value facts — single-value predicates (uses_database / deployment_platform / auth_method) are exclusive claims Claude follows authoritatively, so a stale one is the most dangerous kind of memory. The 0.12 harm benchmark caught Claude emitting git push heroku HEAD:main from a stale deployment_platform fact with zero hedge — and supersession only protects against this if the replacement was recorded. New Recall::StalenessAnnotator (pure function) flags single-value facts that are old (valid_from/created_at older than injection_stale_days, default 180) AND not recently confirmed (last_recalled_at null or stale); Hook::ContextInjector appends a ⚠ stale: recorded YYYY-MM-DD … verify before relying marker at SessionStart so Claude can hedge or verify instead of blindly following. Multi-value predicates are never annotated (they accumulate; one stale entry isn't authoritative). New Configuration#injection_stale_days (CLAUDE_MEMORY_INJECTION_STALE_DAYS), deliberately much longer than the 14-day dashboard review window. Serves the 1.0 long-horizon-quality pillar — it's the first defense against memory degrading session quality over months.
  • Negative-fact harm benchmark — full 13-scenario corpus + release gate — expands the 0.11 3-scenario prototype to 13 cases across four harm classes (stale_tech, mismatched_scope, superseded_undetected, and the new reference_material_as_fact). Each scenario ships a project_files scaffold whose current state contradicts the wrong memory fact, so the test measures "does Claude follow stale/wrong memory over the project's actual state?" rather than reacting to an empty directory. Scored best-of-N (default 3 runs, majority vote per scenario via HARM_BENCH_RUNS) to absorb single-shot LLM nondeterminism. HARM_RATE_THRESHOLD (default 1%) fails the run if the majority-harmed scenario rate is exceeded — making "memory doesn't make Claude wrong" a measurable release gate rather than a marketing claim. The first full-corpus real-mode run surfaced a real harm (stale deployment fact) and a harness confound (empty-tmpdir noise), which drove both the staleness guard above and the scaffold + best-of-N harness hardening.
  • claude-memory audit — memory health diagnostic — productionizes the 2026-05-21 contamination audit into a stable diagnostic surface anyone using claude_memory can run on their own setup. Ten contract checks (C001-C010) cover open conflicts, single-cardinality multiplicity, distillation backlog, shortcut-leak detection, duplicate global conventions, bare-conclusion rate, project starvation, auto-memory import gaps, and single-cardinality churn. --json is the stable contract for CI; --severity filters; --no-exit always exits 0. The /audit-memory slash command wraps the same runner for an interactive walkthrough. docs/audit_runbook.md documents each check's rationale and remediation. CHECK_METHODS is append-only by design so JSON consumers don't break when new checks land. New claude-memory import-auto-memory retroactively pulls ~/.claude/projects/<slug>/memory/*.md entries that AutoMemoryMirror previously missed (slug bug: tr("/", "-") left underscores intact, so claude_memory paths never matched). Contributes to the visibility pillar of 1.0.
  • Contamination guardrails — ReferenceMaterialDetector example-quote guard + Resolver :discard path — the distiller used to treat example sentences in docs/CLAUDE.md ("e.g., postgres", "for example, mysql") as literal claims about the project, accumulating 103 rejected single-cardinality facts over six weeks before being caught by the 2026-05-21 audit. Two defenses now: (1) ReferenceMaterialDetector flags single-cardinality predicate extractions whose source text contains e.g., / for example / i.e. quote patterns so they're tagged reference material at write time; (2) Resolver gains a :discard resolution path for the same shape so the fact never lands even if the detector misses. Memory shortcuts (memory.decisions / .conventions / .architecture) refactored from FTS text search (which returned facts whose object matched the predicate keyword) to predicate-based filtering via PredicatePolicy, with project-DB precedence over global. Closes a class of "is memory still trustworthy?" bugs that erode the 1.0 stability claim.
  • OpenTelemetry ingestion + dashboard Telemetry tab — Claude Code can now export metrics, log-style events, and (opt-in) traces straight into the dashboard via OTLP/HTTP/JSON. New claude-memory otel CLI manages the env block in .claude/settings.json (--enable, --disable, --enable-traces, --capture-prompts, --status, --verify); the dashboard exposes /v1/metrics, /v1/logs, /v1/traces on 127.0.0.1:3377 and a new "Telemetry" drawer showing cost per hour, tokens by model, top tools by latency, and a per-prompt journey waterfall that UNIONs otel_events with the existing activity_events. Schema v18 adds otel_metrics/otel_events/otel_traces plus an additive prompt_id column on activity_events for journey correlation. Privacy posture: nothing past metric counts is captured by default; OTEL_LOG_USER_PROMPTS only flips on with explicit --capture-prompts confirmation; traces remain 501-gated until the user opts in. Sweep retention defaults: 30 days metrics, 14 days events, 7 days traces.
  • Pre-release hook smoke gate (bin/pre-release-smoke) — verifies the installed claude-memory gem actually fires hooks correctly and populates expected detail_json fields per spec/smoke/expected_fields.yml. Codifies the verification convention from feedback_hooks_run_installed_gem.md into a machine-enforced release gate. The trap has been sprung twice (2026-04-16 ActivityLog, 2026-04-30 #47 token-budget); the gate exists so it can't be sprung a third time. Wired into the /release skill as Phase 1 Step 6 (after specs, before lint). First 0.12.0 milestone item.
  • /study-repo memory-discipline guard (prompt-only) — top-level "CRITICAL: Memory Discipline" section in .claude/skills/study-repo/SKILL.md explicitly forbids the LLM from extracting external projects' tech stack as project-level facts. Roots the cleanup work claude-memory reject had to do during 0.11 (27-fact misattribution cluster on 2026-04-23/24, see quality_review.md 2026-04-30 cause-4 finding). Defense-in-depth detector deferred to 0.12.x or later, only built if measurement shows persistent leakage.
  • API stability audit (docs/api_stability.md) — authoritative public-API contract enumerating which CLI commands, MCP tools, hook events, Ruby classes, and schema surfaces are stable / experimental / internal. Default-to-internal applied throughout; the doc is the source of truth for what 1.0's semver promise will lock down. New ClaudeMemory::Deprecations.warn(name:, replacement:, removed_in:) module wired into PredicatePolicy.canonicalize as the first soft-rename — has_convention and primary_language synonyms now emit deprecation warnings scheduled for removal in 1.0.0. README + CLAUDE.md link to the new doc; suppress noise via CLAUDE_MEMORY_NO_DEPRECATIONS=1.
  • Release-to-release benchmark scoreboard — bin/run-evals now writes spec/benchmarks/results/<version>.json after each run; new bin/bench-diff compares the current scoreboard against the most recent prior tagged version's and exits non-zero if any tracked pass-rate dropped beyond the threshold (default -5%, configurable via --threshold). Wired into /release skill Phase 1 as Step 7 — the release aborts on regressions before publish. First release with this gate is 0.12.0 itself; from 0.13.0 onward bench-diff actively gates against 0.12 baselines.

Deferred to 0.13

  • CLAUDE.md comparative baseline numbers (#4) — the comparative E2E harness compares static CLAUDE.md (auto-loaded into context) against ClaudeMemory's MCP-tool retrieval, but in headless claude -p mode Claude doesn't proactively call the recall tools, so the comparison doesn't yet exercise ClaudeMemory's retrieval path fairly (first run returned a misleading ClaudeMemory 0/10 = no-memory 0/10 vs CLAUDE.md 8/10). Publishing that would mislead, so the numbers are withheld and the harness fix is tracked for 0.13. This surfaced a genuine separable observation — in fully headless, non-tool-forcing usage, ClaudeMemory's contribution rides entirely on the SessionStart context-hook injection — also tracked for 0.13. See docs/1_0_punchlist.md #4 / #16.

Upgrade Notes

  • Schema migrates automatically to v18 (OTel telemetry tables + prompt_id on activity_events) on first DB open via Sequel::Migrator — no manual step. Round-trip migration specs cover the upgrade path from prior release boundaries.
  • The staleness marker now appears in SessionStart context for single-value facts (uses_database / deployment_platform / auth_method) older than 180 days and not recently recalled. This is additive and advisory (a ⚠ stale … verify before relying note). Tune the window with CLAUDE_MEMORY_INJECTION_STALE_DAYS; the existing CLAUDE_MEMORY_STALE_DAYS (dashboard review window) is unchanged.
  • No breaking API changes. has_convention / primary_language predicate synonyms continue to emit deprecation warnings (scheduled for removal in 1.0.0); suppress via CLAUDE_MEMORY_NO_DEPRECATIONS=1.

🧪 Real Eval Validation

Results: 2/6 passed ⚠️ 4 failed
Duration: 101.46s
Estimated Cost: ~$0.12

⚠️ Some real eval tests failed. Check the workflow logs for details.