Skip to content

v0.13.1 — Observational Layer: Audited & Repaired

Choose a tag to compare

@codenamev codenamev released this 24 Jun 11:41
· 65 commits to main since this release
Immutable release. Only release title and notes can be modified.

Theme: The observational layer, audited and repaired. A critical examination of every observation in a real (dogfooding) project DB found the episodic layer was producing ~no useful observations and injecting noise into sessions — three evidence-backed defects, now fixed (the data is in docs/improvements.md #72–#75). No schema changes, no breaking changes.

Fixed

  • High-precision Layer-1 observation filter (#74). The high-recall Observer was scraping code, doc, and transcript fragments past noise_body? and injecting them into SessionStart (measured: 38 of 117 obvious-noise rows slipped through — spec fixtures, CHANGELOG table rows, benchmark tree output, even the distiller's own source comments). Strengthened the gate to reject code/JSON key: "value", method calls, table pipes, box-drawing glyphs, (vector) labels, and raw JSONL fields, and to require a body to begin like a prose sentence. Verified against the real noise corpus: every sampled fragment now rejected, clean prose decisions/conventions kept.
  • Corroboration can finally accumulate (#73). Dedup matched on exact normalized strings, so varied wording of the same event never folded — every observation stayed at corroboration_count = 1 and the promotion gate could never fire. Replaced exact grouping with greedy clustering over an injected similarity matcher; the default Observe::TokenOverlapMatcher (lexical Jaccard, deterministic, free) folds near-duplicates so corroboration climbs toward promotion. Pure synonym paraphrases still need real embeddings (measured: tfidf can't separate them from unrelated text on short bodies) — injectable via the matcher: seam.
  • Observation capture elevated to a first-class SessionStart ask (#72). Authoring observations was a paragraph buried in the optional deep-distill prompt, which fires almost never (store_extraction had zero calls in the layer's lifetime; Layer-1 auto-ingest carried ~100:1 of the load). Decoupled it into its own prominent ## Log What Happened section. Whether the LLM-authored (Layer-2) path now fires is measurable via the mcp_extraction content-item source.

🧪 Real Eval Validation

Results: 4/6 passed ⚠️ 2 failed
Duration: 104.16s
Estimated Cost: ~$0.12

⚠️ Some real eval tests failed. Check the workflow logs for details.