Skip to content

v0.59.18.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 20:32
f7c8ccd

Dream no longer keeps made-up quotes, wrong-speaker quotes or invented numbers as memory, and gbrain eval compare computes the statistics it claims.

When the nightly dream cycle turns a conversation into brain pages, a mechanical check compares each quote with the transcript. Until now, a quote it could not find lost its quotation marks and stayed on the page as ordinary text, so an invented sentence became searchable memory. Pages that already existed (person pages, earlier reflections) were not checked at all, a close-match repair could splice in the next speaker's words, and numbers the transcript never mentioned were only counted.

Now any sentence that fails the check leaves the page body and is kept word for word in the page's unverified_claims frontmatter, which get_page shows but search, recall and think do not read. A sentence fails when its quote appears in no source transcript, only matches across two speakers, is attributed to someone other than the person who said it, or when it states a number or date the transcript does not contain. Pages that already existed are checked on the sentences the run added, against the transcripts that wrote them. Timeline entries, facts and links derived from a failed sentence are removed with it. Each kept quote records its source file, character span and speaker in grounding.quotes.

Measured on a fixed three-cycle experiment (3 transcripts, a scripted model that writes 13 supported claims and 17 invented ones, including edits into existing pages), counting what search and recall can read after the third cycle:

After 3 dream cycles v0.59.13.0 v0.59.18.0
Invented claims in active memory (of 17) 17 2
of which fabricated quotes (of 6) 6 0
of which speaker-swapped quotes (of 3) 3 0
of which invented numbers (of 6) 6 0
of which unquoted inventions with no number (of 2) 2 2
Supported claims kept (of 13) 13 13

The fixture was written alongside the check, so treat it as a regression pin, not a hallucination rate for a real model. Plain-prose inventions with no quote or number are not mechanically checkable and still get through.

gbrain eval compare printed "paired bootstrap with Bonferroni correction" but computed only side-by-side averages. It now computes paired statistics from per-question rows: for runs whose ledger record points at their per-question output, it joins the rows by question and reports a 95% bootstrap interval, a p-value and a Holm-corrected p-value for each metric. Runs without per-question rows are labeled aggregate-only, with no significance claim.

To take advantage of v0.59.18.0

Run gbrain upgrade. There is no migration. The check applies to the next dream cycle; existing pages are not rescanned. To review what was held back, run gbrain get <slug> and read unverified_claims. To compare two LongMemEval runs with real statistics, record both with gbrain eval longmemeval --record --output <file> and run gbrain eval compare --baseline <run_id> --candidate <run_id>.

Itemized changes

  • synthesize-verify.ts checks sentence-sized claim units (sentences, list items, table rows) instead of bare quote spans, and quarantines a failing unit whole instead of removing its quotation marks.
  • Close-match quote repairs stay inside one speaker's turn and are trimmed to the matched words (write-path audit C-6); any match that touches a speaker label is refused.
  • Attribution check: when a sentence names a transcript speaker, the quote must come from that speaker's turn.
  • Numbers and dates are compared by value, so $250K, $250,000 and 250 thousand agree, as do 2026-03-14 and March 14th; the transcript file name counts as a source for dates.
  • Verification covers every page a child wrote. A page created during the run is checked whole; an older page is checked on the sentences missing from its revision before the run (page_versions), against every transcript that wrote it. Epochs come from the child jobs' creation time, so resumed children still count their own pages.
  • The unmanaged write-back re-projects timeline entries, facts, takes and links from the verified body in the same transaction. The managed path already re-projected timeline entries, facts and takes through its page publication.
  • New telemetry in details.synthesis.quote_verify: quarantined_claims, pages_with_quarantine, preexisting_diffed, skipped_unchanged and one counter per failure reason. stripped, skipped_preexisting and numeric_claim_warns are gone.
  • gbrain eval compare: --baseline, --candidate, --draws, --seed; JSON gains paired and paired_unavailable; Markdown gains a "Paired comparisons" table with a significant / not significant verdict. A per-question file path in the ledger must resolve inside the repository root; paths that escape it are refused and never read. The statistics module (src/core/eval/paired-bootstrap.ts) is a port of the gbrain-evals situation-recall comparator with a two-sided p-value.

For contributors

  • New tests: test/cycle-repeated-consolidation.test.ts (fails on v0.59.13.0, passes here), test/eval-paired-bootstrap.test.ts; extended test/cycle-synthesize-verify.test.ts, test/eval-compare.test.ts, test/cycle-write-path-mini-eval.test.ts.
  • bun run scripts/repeated-consolidation-experiment.ts [--per-cycle] prints the experiment as JSON ($0, no network). Record in docs/eval/FIX_WAVE_BASELINES.md.