v0.12.0 — Release Discipline, Observability, Self-Audit
·
117 commits
to main
since this release
Immutable
release. Only release title and notes can be modified.
Theme: Release Discipline + Observability + Self-Audit — the infrastructure that makes a 1.0 semver promise defensible. This release locks down the public API surface, adds the observability primitives (OTel ingestion, dashboard Telemetry) and the self-audit toolkit (claude-memory audit) that serve the visibility pillar, and ships the negative-fact harm benchmark + staleness guard that make the long-horizon-quality claim measurable rather than aspirational.
Added
- Staleness guard for single-value facts — single-value predicates (
uses_database/deployment_platform/auth_method) are exclusive claims Claude follows authoritatively, so a stale one is the most dangerous kind of memory. The 0.12 harm benchmark caught Claude emittinggit push heroku HEAD:mainfrom a staledeployment_platformfact with zero hedge — and supersession only protects against this if the replacement was recorded. NewRecall::StalenessAnnotator(pure function) flags single-value facts that are old (valid_from/created_atolder thaninjection_stale_days, default 180) AND not recently confirmed (last_recalled_atnull or stale);Hook::ContextInjectorappends a⚠ stale: recorded YYYY-MM-DD … verify before relyingmarker at SessionStart so Claude can hedge or verify instead of blindly following. Multi-value predicates are never annotated (they accumulate; one stale entry isn't authoritative). NewConfiguration#injection_stale_days(CLAUDE_MEMORY_INJECTION_STALE_DAYS), deliberately much longer than the 14-day dashboard review window. Serves the 1.0 long-horizon-quality pillar — it's the first defense against memory degrading session quality over months. - Negative-fact harm benchmark — full 13-scenario corpus + release gate — expands the 0.11 3-scenario prototype to 13 cases across four harm classes (stale_tech, mismatched_scope, superseded_undetected, and the new reference_material_as_fact). Each scenario ships a
project_filesscaffold whose current state contradicts the wrong memory fact, so the test measures "does Claude follow stale/wrong memory over the project's actual state?" rather than reacting to an empty directory. Scored best-of-N (default 3 runs, majority vote per scenario viaHARM_BENCH_RUNS) to absorb single-shot LLM nondeterminism.HARM_RATE_THRESHOLD(default 1%) fails the run if the majority-harmed scenario rate is exceeded — making "memory doesn't make Claude wrong" a measurable release gate rather than a marketing claim. The first full-corpus real-mode run surfaced a real harm (stale deployment fact) and a harness confound (empty-tmpdir noise), which drove both the staleness guard above and the scaffold + best-of-N harness hardening. claude-memory audit— memory health diagnostic — productionizes the 2026-05-21 contamination audit into a stable diagnostic surface anyone using claude_memory can run on their own setup. Ten contract checks (C001-C010) cover open conflicts, single-cardinality multiplicity, distillation backlog, shortcut-leak detection, duplicate global conventions, bare-conclusion rate, project starvation, auto-memory import gaps, and single-cardinality churn.--jsonis the stable contract for CI;--severityfilters;--no-exitalways exits 0. The/audit-memoryslash command wraps the same runner for an interactive walkthrough.docs/audit_runbook.mddocuments each check's rationale and remediation.CHECK_METHODSis append-only by design so JSON consumers don't break when new checks land. Newclaude-memory import-auto-memoryretroactively pulls~/.claude/projects/<slug>/memory/*.mdentries thatAutoMemoryMirrorpreviously missed (slug bug:tr("/", "-")left underscores intact, soclaude_memorypaths never matched). Contributes to the visibility pillar of 1.0.- Contamination guardrails —
ReferenceMaterialDetectorexample-quote guard +Resolver:discardpath — the distiller used to treat example sentences in docs/CLAUDE.md ("e.g., postgres", "for example, mysql") as literal claims about the project, accumulating 103 rejected single-cardinality facts over six weeks before being caught by the 2026-05-21 audit. Two defenses now: (1)ReferenceMaterialDetectorflags single-cardinality predicate extractions whose source text containse.g.,/for example/i.e.quote patterns so they're tagged reference material at write time; (2)Resolvergains a:discardresolution path for the same shape so the fact never lands even if the detector misses. Memory shortcuts (memory.decisions/.conventions/.architecture) refactored from FTS text search (which returned facts whose object matched the predicate keyword) to predicate-based filtering viaPredicatePolicy, with project-DB precedence over global. Closes a class of "is memory still trustworthy?" bugs that erode the 1.0 stability claim. - OpenTelemetry ingestion + dashboard Telemetry tab — Claude Code can now export metrics, log-style events, and (opt-in) traces straight into the dashboard via OTLP/HTTP/JSON. New
claude-memory otelCLI manages the env block in.claude/settings.json(--enable,--disable,--enable-traces,--capture-prompts,--status,--verify); the dashboard exposes/v1/metrics,/v1/logs,/v1/traceson127.0.0.1:3377and a new "Telemetry" drawer showing cost per hour, tokens by model, top tools by latency, and a per-prompt journey waterfall that UNIONsotel_eventswith the existingactivity_events. Schema v18 addsotel_metrics/otel_events/otel_tracesplus an additiveprompt_idcolumn onactivity_eventsfor journey correlation. Privacy posture: nothing past metric counts is captured by default;OTEL_LOG_USER_PROMPTSonly flips on with explicit--capture-promptsconfirmation; traces remain 501-gated until the user opts in. Sweep retention defaults: 30 days metrics, 14 days events, 7 days traces. - Pre-release hook smoke gate (
bin/pre-release-smoke) — verifies the installed claude-memory gem actually fires hooks correctly and populates expecteddetail_jsonfields perspec/smoke/expected_fields.yml. Codifies the verification convention fromfeedback_hooks_run_installed_gem.mdinto a machine-enforced release gate. The trap has been sprung twice (2026-04-16 ActivityLog, 2026-04-30 #47 token-budget); the gate exists so it can't be sprung a third time. Wired into the/releaseskill as Phase 1 Step 6 (after specs, before lint). First 0.12.0 milestone item. /study-repomemory-discipline guard (prompt-only) — top-level "CRITICAL: Memory Discipline" section in.claude/skills/study-repo/SKILL.mdexplicitly forbids the LLM from extracting external projects' tech stack as project-level facts. Roots the cleanup workclaude-memory rejecthad to do during 0.11 (27-fact misattribution cluster on 2026-04-23/24, seequality_review.md2026-04-30 cause-4 finding). Defense-in-depth detector deferred to 0.12.x or later, only built if measurement shows persistent leakage.- API stability audit (
docs/api_stability.md) — authoritative public-API contract enumerating which CLI commands, MCP tools, hook events, Ruby classes, and schema surfaces are stable / experimental / internal. Default-to-internal applied throughout; the doc is the source of truth for what 1.0's semver promise will lock down. NewClaudeMemory::Deprecations.warn(name:, replacement:, removed_in:)module wired intoPredicatePolicy.canonicalizeas the first soft-rename —has_conventionandprimary_languagesynonyms now emit deprecation warnings scheduled for removal in1.0.0. README + CLAUDE.md link to the new doc; suppress noise viaCLAUDE_MEMORY_NO_DEPRECATIONS=1. - Release-to-release benchmark scoreboard —
bin/run-evalsnow writesspec/benchmarks/results/<version>.jsonafter each run; newbin/bench-diffcompares the current scoreboard against the most recent prior tagged version's and exits non-zero if any tracked pass-rate dropped beyond the threshold (default -5%, configurable via--threshold). Wired into/releaseskill Phase 1 as Step 7 — the release aborts on regressions before publish. First release with this gate is 0.12.0 itself; from 0.13.0 onward bench-diff actively gates against 0.12 baselines.
Deferred to 0.13
- CLAUDE.md comparative baseline numbers (#4) — the comparative E2E harness compares static CLAUDE.md (auto-loaded into context) against ClaudeMemory's MCP-tool retrieval, but in headless
claude -pmode Claude doesn't proactively call the recall tools, so the comparison doesn't yet exercise ClaudeMemory's retrieval path fairly (first run returned a misleading ClaudeMemory 0/10 = no-memory 0/10 vs CLAUDE.md 8/10). Publishing that would mislead, so the numbers are withheld and the harness fix is tracked for 0.13. This surfaced a genuine separable observation — in fully headless, non-tool-forcing usage, ClaudeMemory's contribution rides entirely on the SessionStart context-hook injection — also tracked for 0.13. Seedocs/1_0_punchlist.md#4 / #16.
Upgrade Notes
- Schema migrates automatically to v18 (OTel telemetry tables +
prompt_idonactivity_events) on first DB open viaSequel::Migrator— no manual step. Round-trip migration specs cover the upgrade path from prior release boundaries. - The staleness marker now appears in SessionStart context for single-value facts (
uses_database/deployment_platform/auth_method) older than 180 days and not recently recalled. This is additive and advisory (a⚠ stale … verify before relyingnote). Tune the window withCLAUDE_MEMORY_INJECTION_STALE_DAYS; the existingCLAUDE_MEMORY_STALE_DAYS(dashboard review window) is unchanged. - No breaking API changes.
has_convention/primary_languagepredicate synonyms continue to emit deprecation warnings (scheduled for removal in 1.0.0); suppress viaCLAUDE_MEMORY_NO_DEPRECATIONS=1.
🧪 Real Eval Validation
Results: 2/6 passed
Duration: 101.46s
Estimated Cost: ~$0.12