Repository navigation
[0.22.0] — 2026-10-01
Compatibility: project isolation and retired-record withholding are now
the defaults for scoped callers. Once a request's project is resolved, search,
POST /resolve and /context/prime leave out records tagged to a different
project; pass include_other_projects on search or POST /resolve to cross
projects. POST /resolve and search's evidence mode leave out retired records;
pass include_retired=true to see history. An MCP server over HTTP no longer
scopes requests to its own working directory: a remote client sends its project
in the X-Palinode-Project header (palinode mcp-config --http --project <slug>),
and a request with no project is unscoped. The OpenClaw plugin's palinode_search tool and CLI search are now
project-scoped by default. An explicit read by path (palinode_read, GET /read,
palinode read) still works across projects.
Contributors: Tom Tang (PR #227); NesrineGharbi77 (PR #225); Vardhman Gupta (PR #228); Isha Zaka (PR #230).
Added
SECURITY.md: a memory-poisoning and trust limitations statement. What the "memory is data, not instructions" delivery framing, delivery-time withholding of agent-directed text, provenance and project isolation do and do not protect against, with the live-agent evaluation's numbers behind the claims (n=6 per cell): Palinode beats a plain instructions file on forged corrections, authority laundering and cross-project steering, and an embedded instruction still gets followed byclaude-haiku-4-5at roughly plain-file parity (2–3/6) even after the delivery-time fix. States plainly what provenance and corroboration do not guarantee, which isolation gaps this release closed (the hook's deadline fallback, the OpenClaw plugin's recall, and automaticcross_refsnaming other projects' records), and what to do with imported or shared memory. Linked from the README anddocs/HARNESSES.md.docs/HARNESSES.md: coexisting with a client's own memory. Palinode never writes Claude Code's auto-memory or Codex's memories, and reads client files only on request (import, opt-in transcript capture). Measured in the live-agent evaluation: no native memory file changed because of Palinode in any run. The two stores are not reconciled, and agents rarely report disagreements, so keep a fact in one place.- Evaluation wrap-up tools. Optional pinned free-text judge (
bench.agent_tasks judge: a model different from the agent's, retries, resume, and a seeded hand-audit sheet), a house-rule row (row 18), and a row-8 correction that resubmits the untouched sentence when the shipped route refuses a content-losing replacement. bench/agent_tasks: a live-agent task harness. A versioned corpus of small synthetic coding tasks where a recorded decision changes the right action — all 17 rows of the study (rejected approaches, decision changes after restarts, consolidation and transcript capture, cross-session and cross-client continuation, archive/restore, conflicts, embedded instructions, forged corrections, copied claims, cross-project leakage, privileged handoffs, native-memory disagreement), each with its control, plus a derived held-out split — andpython3 -m bench.agent_tasks plan|grade|report.planstamps every cell with its own random tokens and is byte-identical for a given seed;gradescores diffs, answers and delivery records by exact token match, keeps delivered and acted separate, names each failure's first failing stage, and reports missing evidence as NOT RUN rather than a fail;reportprints denominators, measured cost and a full-run projection against the run's cost cap, never a combined score. The harness runs no client itself.palinode doctornames fragmented project tags. A new check,project_tags_unmapped, warns aboutproject/*tags on 10 or more files that are not a canonical ref or a member of a group in the store'sentity-aliases.yaml, and not aproject_maptarget, with their counts. Recall isolation treats those as other projects.- A delivery says whether it found a confident answer. Every search delivery now
carries aconfidenceverdict —confident,weakornone— in
receipt.retrieval, with the per-arm evidence behind it, on every surface. It is
read from each arm's own pre-fusion score (real cosine, normalized BM25), both of
which every delivered row now carries alongside the fused one, because the fused score
is a rank: the top hit is 1.000 whether it is the answer or the least bad of a weak
field, so a caller could not condition on it. The MCP rendering leads withNo confident match.and the arm scores when that is the verdict, and still lists the
weak results underneath; the retrieval log records the verdict on every row of the
delivery, including the single row an empty delivery writes. Measured on the packaged
relevance fixture in both retrieval modes: no no-answer question reaches the confident
mark, and the only answerable questions callednoneare ones whose slate held no
relevant record at all. Marks and their measurements are documented in
palinode/core/confidence.py; the contract is indocs/lexical-retrieval.md
. - In hybrid mode a confident verdict needs both arms. The vector arm may not claim
confidentover a query the keyword arm does not reach its own weak mark on; the
delivery reportsweakwithretrieval.corroboration: "missing"and a diagnostics
line saying why. Measured against the whole fixture corpus with real embeddings, the
vector arm scored three questions with no answer in the corpus at cosine 0.607–0.635 —
above four answerable questions, so no cosine mark separates them — and all three were
queries the keyword arm barely registered. Corroboration leaves 41 of 48 answerable
questions confident against 0 of 14 no-answer ones; the tightest cosine-only mark that
also reaches zero keeps only 36. One-directional and hybrid-only: the keyword arm
needs no corroboration, and ahybrid: falserequest has no second arm to ask
. search.abstain_on_no_confident_match(defaultfalse) withholds the slate on
the MCP surface when the verdict isnone, saying how many results it withheld so an
abstention is never mistaken for an empty store. Off by default deliberately: the
verdict ships as a signal, not a filter, and every other surface returns its rows
either way — the receipt and retrieval log always record what the store delivered, so
the switch changes what is shown and never what is measured. Measured as its own
benchmark arm in every retrieval mode
(bench/results/relevance-no-confident-match-2026-09-22), recall@k and top-1
unchanged throughout: at the floor the MCP surface actually sends it raises correct
abstention on no-answer questions from 0/14 to 2/14 and withholds two slates, both of
them questions with no answer in the corpus; keyword-only, 2/14 to 6/14 with
injections 128/200 → 110/181; at the stricter REST floor it changes nothing at all,
because those slates were already empty. It withheld nothing correct in any mode
.- A memory's own
confidencefrontmatter is rendered beside its match. The
author-stated 0.0–1.0 value has been storable and returnable insidemetadatasince
it was added and was never shown by the surface an agent actually reads, which made
marking a memory half-sure a note to nobody. - A delivery can be explained after the fact.
palinode explain <bundle_id>,GET /explain/{bundle_id}, thepalinode_explainMCP and plugin tools and the local
inspector page at/ui/delivery/{bundle_id}answer "what context did this agent
receive, and why was this memory selected or qualified?" from the reference a
delivery already hands back. Composed entirely from the existing delivery receipt
and retrieval log — no new table, no new file, no change to what is recorded: the
supplied records, the exact revision each was supplied at and whether that source has
changed since, the server-resolved scope and project, the calling surface and whether
the recall was explicit or passive, each record's disposition, the delivery's coverage
qualifiers and its evaluation time. Output is capped with an explicit "N more not
shown" and visibility-filtered: a record the caller may not see is counted, never
named. Each supplied record links to its memory and its Git history, and a clearly
labelled pointer names where a correction is made: thepalinode corrections
preview / apply / undo flow this release adds, or the plain save and supersede routes
. - Every field an explanation cannot answer is reported as
unavailablewith its reason
from a closed vocabulary, never invented and never silently omitted — including the
difference between "no rows carry this reference" and "this surface writes no rows at
all"./context/prime,/resolveand an empty-query search return receipts and write
no retrieval-log rows, so a reference from one of them is available only in the
response that carried it; a lookup that finds nothing reports every candidate cause
with whether it could be checked. Supplied context is shown separately from evidence
that an agent acted on it, of which Palinode records none. The caller's own query
prose stays diagnostics-only, exactly as on the receipt: the MCP and plugin tools
cannot ask for it. Lookup availability, retention, restart behaviour, which surfaces
write rows, and the questions this storage cannot answer are documented in
docs/DELIVERY-RECEIPTS.md. - A delivery reference does not entitle you to the asker's words. The two fields
that describe the person who asked rather than the memory supplied — the query
prose and the delivery's session id — are served only when the API is bound where
nothing off this machine can reach it (the same bind predicate the local inspector
hard-refuses on, one function so the two cannot drift), or when the caller presents
the session that made the delivery. A bundle id is a short digest that travels in
receipts and thatGET /trace/{file}already lists for every delivery a file was
supplied in, so holding one is not authorisation. Anything else gets the ordinary
public view with both fields markedwithheld_diagnostics_onlyand the reason
attached — never a 403, which would answer a question about the bundle the caller
has not earned, and never a message that varies with whether a session id was
recorded.palinode explain --diagnosticskeeps working against your own loopback
API and prints the withheld marker and the reason against a remote one, so it never
quietly shows less than you asked for. - Correction candidates can be mined from harness session transcripts — the moments a
user overturned a decision, rejected an approach and said why, or asked for something
to be remembered. Off by default and empty by default: the source reads nothing until
capture.transcripts.enabledis set andcapture.transcripts.harness_pathsnames
directories to read (Claude Code only in this release; Codex CLI is a follow-up). A
deterministic narrow grep finds candidate spans in the user's own turns, and a model
is used only to classify a bounded window around a span into a closed set; anything
not confidently one of the three classes is left for review. Asking a model at all is
a third, separate opt-in (capture.transcripts.classify, also off by default): with
it off, a scan runs the deterministic stage only, makes no request to any model
endpoint, and queues every candidate asneeds_reviewwith a provenance that
distinguishes "classifier not run" from "endpoint unreachable" and from "answer
unusable". Whether classification runs is a configuration decision — no CLI flag,
request field or tool parameter can enable it for a single call. Nothing is applied —
every candidate is a proposal in an operational JSONL queue under.palinode/, and no
memory is written, changed or retired by this path. Candidates are deduplicated on
(session id, span hash), so re-reading a transcript or reading a later summary that
restates the same correction creates no second candidate and no corroboration count
. palinode correctionslists the candidate queue with--project,--sinceand
--scan, TTY-aware like the other commands, with the same capability onPOST /corrections, thepalinode_correctionsMCP tool and the plugin. Running a fresh
detection pass (--scan/"scan": true) is deliberately an operator action on the
CLI and REST API; the MCP and plugin tools are read-only listings.- Transcript correction mining discloses what it sends, not only what it stores, and
every surface's report states whether classification ran — naming the model and
endpoint role when it did, and sayingdetection only; nothing was sent to a model
when it did not. Whencapture.transcripts.classifyis on, classifying a candidate
transmits a bounded window of conversation text (at most five turns, 400 characters
each) to the configured consolidation model endpoint, which may be remote. Only the user's own turns and ordinary assistant replies contribute text,
with quoted, pasted and harness-injected regions stripped; every other kind of turn —
tool output, subagent turns, summaries, harness metadata — is replaced by a
placeholder naming its kind, so file contents and credentials in tool output are never
transmitted. The deterministic detection pass sends nothing.palinode controls statusandGET /statusname the destination. - Transcript correction mining honours the existing capture controls exactly as other
automatic capture sources do:capture_pausedstops a scan before the first read, an
excluded project or path skips those transcripts, and the lookback window and
candidate cap log and report what they skipped rather than truncating silently.
palinode controls statusand the/statusdisclosure name the source, what it reads
and what it stores whenever it is enabled. - A generated MCP client can pin its project, so every recall call from that client resolves the same one without passing a project argument:
palinode init --pin-projectwrites the setting into the.mcp.jsonit generates, alongside the existingpalinode mcp-config --stdio --project <slug>. Without the flag the emitted block is unchanged. A pinned value that is not a slug or aproject/<slug>ref is refused rather than quietly replaced by repository inference — the CLI aborts, the MCP tool returns an error and the API answers 400. - Every response that carries project scope now names the project and the source that decided it: search leads with
Scope: project/<slug> (<source>)on the MCP and CLI surfaces, returnsprojectandproject_resolved_byin the REST receipt envelope (the bare result array is unchanged), and the OpenClaw plugin renders the same line. What is reported is what was applied. The session-start digest already reported both and keeps its wording. - An API search request that carries an explicit context — an empty list included — keeps the scope the caller decided on, so an empty one means "no project scope" and is how a caller opts out on a server with a pinned project.
palinode search --no-contextnow sends exactly that instead of omitting the field, which on a pinned server would have silently reinstated the scope it opted out of. bench/relevance/— a versioned, sanitized relevance and abstention fixture (39 synthetic records, 62 labelled questions across current release state, rejected approach, changed decision and no-answer/unrelated-project, with held-out paraphrases and exact-identifier variants) and a runner that scores both the keyword-only and the hybrid arm on a real store: relevant-hit recall@k, irrelevant injections, footer-only hits, redundant results, project isolation, abstention versus confident match, plus useful-context tokens and p50/p95 latency through the payload the client actually receives. The hybrid arm is marked NOT RUN rather than filled with synthetic vectors when no embedding endpoint is reachable. Production search defaults are unchanged; the baseline is recorded inbench/results/.- Benchmark: a labelled prospective-trigger phrasing measurement (
python -m bench.trigger_phrasing) over a versioned synthetic corpus — third-person versus user-worded versus multi-phrasing descriptions, scored with the real embedder and the real trigger match path, reporting separation (AUC), best-threshold precision/recall and fire / false-fire rates for absolute, relative-margin and vector-or-BM25 decision rules. It changes no default and refuses to run without an embedding endpoint. - Reviewing and applying a correction is now one two-phase operation on every surface.
palinode corrections previewreads and writes nothing: it returns the old text, the proposed new text, the affected document (and claim), its exact source revision, the rationale, the correction's source, the project scope, thesupersedes/superseded_byrelation that would be recorded, every other record that quotes or derives from the target — reported, never rewritten — and the exact recovery command.palinode corrections applyis the only writer and needs both the preview's revision and explicit confirmation; a target that changed in between is refused with both revisions, and a ref or claim id matching more than one record is refused with every match rather than resolved by guesswork. It persists only through the existing validated path (the same save, thenarchive_memorywithsuperseded_by), so the git commit, the-history.mdaudit sibling and the index propagation are unchanged; the commit subject and history line name the actor as a reviewed correction and carry the candidate id when one was the source.update_policyis inherited from the target, never imposed. The same four phases ship asPOST /corrections/{preview,apply,dismiss,undo}, thepalinode_correction_preview/_apply/_dismiss/_undoMCP tools (non-core; preview isreadOnlyHint, apply isdestructiveHint) and the matching plugin tools. - Transcript correction candidates are reviewable:
palinode corrections dismiss --candidate <id> --reason …records a reviewer's decision, and--candidate <id>on preview/apply carries the span, session and turn into the correction's provenance.needs_reviewand unclassified candidates review like any other; a candidate whose span never named what it replaced has no target inferred — the reviewer names one. Applied and dismissed rows are marked and kept in the queue, never deleted, which is what stops a re-scan re-proposing the same span: the dedupe key is (session id, span hash), read from every row. When the originating transcript is gone the preview says the source is unavailable and why that does not weaken the stored span. - A documented recovery path with its own preview.
palinode corrections undopreviews by default and writes only with--confirmand the preview's revision. Its output distinguishes three things that are routinely conflated: restoring a previous assertion (what it does), deleting history (not offered — the commits, the history sibling and the replacement all remain), and undoing an agent's external actions (impossible, and stated as such). It refuses to resurrect a record separately retracted or withdrawn by a forget request, pointing atpalinode unretract/palinode forget-withdrawinstead. - The inspector's memory page gained a read-only Correct or retire this section: the record's current revision, what else references it, any pending correction candidates quoting text in it, and the exact copy-pasteable CLI commands for preview, apply and undo. No form, nothing posts, no in-browser mutation.
docs/CORRECTIONS.md— the correction / retirement / recovery walkthrough, linked from the Quickstart, the inspector guide and the data-lifecycle page, which it is careful not to be confused with.archive,restore,unretractandforget-withdrawcan each be previewed.--dry-runon the CLI,dry_run: trueonPOST /archive,/restore,/unretractand/forget-withdraw, the same parameter on the four MCP tools, and adryRunargument on the plugin core's lifecycle client (which gainsarchiveMemorybeside its existing reversals). A dry run validates exactly as an apply would and writes nothing — no file, no history line, no index row, no commit — and shows the record, the frontmatter delta, the relation it would record or remove, the retained copies, and the recovery command.forget-withdrawpreviews each step with that step's own dry run. Apply stays the default, so existing scripts and agents behave exactly as before; the MCP tool count is unchanged.consolidation.nightly.max_prompts_per_project(default 4): how many resumed prompts one nightly pass may send a project whose new notes do not fit one prompt.palinode aliasescurates the store'sentity-aliases.yaml.listshows each group with every ref's indexed file count;add <canonical> <member>…creates or extends a group;remove <member>drops one (an emptied group goes);checkruns the alias lint, marking clusters a group already covers, next to theproject_tags_unmappeddoctor check. Writes apply by default and--dry-runshows the file diff. A ref already in another group is refused without--move, with project refs compared case-insensitively as isolation does; refs must becategory/name. The file is always<PALINODE_DIR>/entity-aliases.yaml, written sorted and committed in the store's git, and no memory file is touched. The same operations areGET /aliases,GET /aliases/check,POST /aliases/addandPOST /aliases/remove(withdry_run). There is deliberately no MCP tool: alias groups decide what project-scoped recall shows, so editing them stays with the operator. A write is seen by the next lookup at once, in the writing process and in any other process that reads the file.docs/ENTITY-ALIASES.md— the alias file's format and home, query-time resolution, the canonical ref, case-insensitive project comparison, which spellings to merge and which to keep apart, and how it composes withcontext.project_mapand project isolation. Linked from HARNESSES.md, HOW-MEMORY-WORKS.md and theproject_tags_unmappedremediation, which now namespalinode aliases add.
Changed
- A search returns what the store has evidence for, not
limitresults regardless. Two newsearch.*settings bound the delivered slate on every surface, both relative to the best candidate in the same result set — so the top match always survives and a query with any match still gets one.lexical_fts_threshold(default 0.35) is the keyword arm's floor when it is the only arm (retrieval_mode: lexical, and the keyword fallback); that path used to hard-set the floor to 0.0, which is what filledlimitunconditionally.max_chunks_per_file(default 1) stops one file taking several slots while other files are still competitive; the overflow is deferred rather than dropped, so it still fills a slate that would otherwise come back short and the cap never shortens a delivery. Set them to0.0and0for the previous behaviour. Measured on the relevance rig (keyword arm, 62 questions,top_k5, real SQLite + FTS5): results delivered 268 → 200, irrelevant injections 197/268 → 128/200, repeat slots for one file 10/268 → 2/200, useful-context tokens 19.3% → 26.3%, payload 502.6 → 377.2 tokens per question, while relevant-hit recall rose 48/52 → 49/52 and no question regressed on any counted outcome. The delivery receipt and the retrieval log follow what was delivered, so a shorter slate means fewer rows in both. Abstention is deliberately unchanged: a relative floor cannot return nothing. - The vector arm is now bounded against its own best match, so hybrid search stops padding the slate.
search.vector_relative_floor(new, default 0.85) keeps a vector candidate only when its cosine is at least that fraction of the best cosine in the same candidate set — the counterpart tofts_thresholdon the keyword arm, applied before fusion on the arm's own scores. The absolute cosine floors are unchanged and unmoved (mcp_threshold0.40,api_threshold0.50): an absolute floor says whether a candidate is plausible at all, never whether it is plausible beside the match the query actually found, and it was the second question that decided whether the rest oflimitfilled with weak neighbours. Measured against the best candidate, so the top match always survives — this bounds a slate and cannot empty one; abstention remains a separate decision. Set it to0.0for the previous behaviour. 0.85 is calibrated from every result the relevance rig delivered on both hybrid arms with its raw cosine recorded (62 questions,top_k5, realbge-m3): the weakest relevant result in that set sits at 0.884 of the best cosine in its own slate, 0.88 clears it by 0.004 and 0.90 loses two relevant results. Measured on the rig's hybrid arms — the shipped default mode — at the shipped cosine floor 0.40: results delivered 304 → 262, irrelevant injections 224/304 (73.7%) → 186/262 (71.0%), redundant results 15 → 11, useful-context tokens 20.1% → 23.5%, payload 502.0 → 429.0 tokens per question, with relevant-hit recall, answerable questions with a hit and top-1 all unchanged (52/52, 48/48, 37/48) and the useful-token count identical on both sides — the payload lost exactly the part of it that was not the answer. At the 0.50 floor: 277 → 240 delivered, 198 → 164 injections, 21.5% → 25.1% useful, 469.4 → 403.1 tokens. No question regressed on relevant hits found, rank of the first relevant hit or top-1, and correct abstention is unchanged in both arms. - The keyword arm's score now means the same thing on any store.
search_ftsnormalized BM25 as|bm25| / 25, an absolute constant divided into a quantity that grows with the corpus (IDF ≈ log N/df) and with the number of query terms — so the same exact-identifier hit scored 0.021 on a three-record store, 0.107 on a twenty-record one and 0.204 on a two-hundred-record one, and every mark read against it was pessimistic on a small store. The divisor is now what that query could score in that index: the sum of its units' IDFs, which is the BM25 of a chunk holding every query term once at average length.1.0is that chunk, a denser one reads above it, and the same identifier now scores 1.031 / 1.043 / 1.045 on those three stores. A unit the store has never seen is priced as the rarest a term can be, so a query full of words the store does not know cannot be trivially "covered" — the shape of a question with no answer in the corpus. Nothing about which results are delivered changes: every candidate of one query is divided by the same positive number, so each candidate's ratio against the best one is identical and both relative floors (search.fts_threshold0.40,search.lexical_fts_threshold0.35) read exactly what they read before, with RRF fusing ranks either way. Measured end to end on the relevance rig, keyword arm, before and after: results delivered 200 = 200, recall@k 94.2% = 94.2%, top-1 32/48 = 32/48, irrelevant injections 128/200 = 128/200, payload 377.2 = 377.2 tokens per question — identical, metric for metric, with p50 latency 1.82 → 1.89 ms for the onecount(*)per query unit this costs. - The keyword confidence marks moved with the scale under them: confident
0.22→0.60, weak0.13→0.22, re-calibrated on the packaged relevance fixture with the same discipline as before. 0.60 is the middle of a plateau — 33 answerable questions clear it at every value from 0.56 to 0.62, each with a relevant record in its slate, and the highest a no-answer question reaches is 0.557. 0.22 is the last value that calls no answerable questionnone. On the keyword-only arm, confident answers go 26 → 33 with no-answer confidents still 0, correct abstention 2/14 → 8/14 with the abstain switch on, and all eight answerable exact-identifier questions move fromweaktoconfident— the defect the rescale exists to fix. The vector marks (0.60 / 0.50) and the corroboration rule are untouched; that the two arms' confident marks now share a number is a coincidence of two calibrations, not a shared scale. Measured on the hybrid arms against a real embedder, at both cosine floors and with the abstain switch off or on: answerable confidents 41 → 46, every one with a relevant record in its slate, no-answer confidents still 0, and recall@k, top-1 and correct abstention unchanged. - Search-threshold documentation now identifies the
bge-m3calibration of the MCP and API floors and points to the abstention benchmark for re-checking them with another embedding model. - The nightly consolidation pass selects notes by a per-project watermark instead of a UTC calendar-day window. Each project carries the timestamp of the last pass that resolved it, and the next pass reads the notes written since — so a note reaches the model once, a failed or deferred pass re-sends only what it failed on, and a pass where one project fails does not re-send the projects that succeeded. The marks live beside the activity gate's clock in
<memory_dir>/.palinode/consolidation-state.jsonunderwatermarks; they are derived operational state, never memory content, and survive a restart. Upgrading: nothing to do. A project with no mark cold-starts from the last successful nightly the store recorded — therunshistory, or the activity gate's clock — clamped to the catch-up bound; a store that records no successful pass starts one bound back. The weekly pass keeps its window and is unchanged. consolidation.nightly.lookback_days, and--days Non the nightly cron line, now mean the catch-up bound rather than a lookback window — how far back a cold or long-failed watermark may reach, so one abandoned project cannot hand the model months of notes in a single request. The default moves from 1 to 7 to matchconsolidation.auto_gate.max_hours_elapsed(168 h): the gate may let a week pass before firing a pass at its ceiling, and a shorter bound would drop the notes the gate itself chose to wait on. Existing cron lines keep working and cover at least what they used to; the flag is not removed. The nightly names the bound and its new meaning in its log on every run, and when the bound clamps a mark it reports the project, the mark, the bound and how many note files fell in the gap — never a silent truncation.docs/DATA-LIFECYCLE.mdnow separates the four operations people conflate — correcting a record, archiving/forgetting it (stop using it by default), restoring it, and physically erasing it — with the command for each, a one-paragraph lifecycle summary other pages can link to, and an explicit list of what Palinode cannot erase: provider and client conversation history, context already supplied to a running agent, exports a recipient holds, and clones or backups you do not control. The "where a memory lives" table now says, per location, whether removal is manual or unsupported and how to check absence — every location takes an operator step, and the one that comes closest to free (the SQLite/FTS/vector index) rebuilds its contents automatically but still needs its files deleted. The erasure runbook gains the three steps whose absence made it incomplete: rewrite commit messages as well as blobs (Palinode writes the memory's path into every commit subject, so a path-only rewrite leaves a history that says who the record was about); replace a remote repository rather than force-pushing to it (a force-push moves refs and leaves every old object behind); and delete.palinode.dbwith its-waland-shmsidecars rather than reindexing over it — a reindex drops rows but does not scrub the pages they occupied, and under WAL a committed row is readable in the sidecar until a checkpoint folds it in, so a backup or crash taken while the daemon holds the database open captures the erased text after every row that held it is gone. Verification is by absence: a byte-level grep of the raw database files (a search queries exactly the rows that are already gone) and a sweep that reads git objects rather than grepping a.gitdirectory full of compressed packfiles.palinode doctor'sconsolidation_last_runthresholds now follow what a nightly failure actually costs. A failed pass no longer drops a day of notes, because it does not advance the watermark, so the error fires at a streak as long as the catch-up bound — the streak that does push a mark past it — and the warning at two consecutive failures, which is "the nightly has stopped working" rather than "notes are about to be lost". On a default install that is warn at two, error at seven; the message states the arithmetic in the numbers the pass actually ran.palinode initnow says so when a project ends up with the memory block in both.claude/CLAUDE.mdandAGENTS.md: Claude Code's documented default is to read itsCLAUDE.mdfiles instead ofAGENTS.mdwhenever aCLAUDE.md,.claude/CLAUDE.mdorCLAUDE.local.mdsits in the working directory or above it, while Codex and otherAGENTS.md-aware harnesses readAGENTS.mdand notCLAUDE.md— so each block is live for one side and inert for the other, and an edit to one does not reach the other. The note also states that a re-run does not reconcile drift between the two blocks (initis idempotent per file, and--forceappends a fresh block rather than replacing an edited one).--dry-runprints the same note. Which filesinitwrites, and when, is unchanged; HARNESSES.md carries the rule.
Fixed
- Automatic cross-references respect project isolation. Generation includes only the record’s own projects and global targets. Reads filter existing automatic links against the reader’s scope in both raw YAML and parsed frontmatter, including MCP, CLI and plugin delivery. Automatic plugin reads (core and trigger delivery) withhold another project's record; an explicit read by path stays allowed, as explicit search can cross projects. Authored body links and typed relations stay unchanged.
- Rapid saves no longer leave a written file uncommitted. A burst of saves could fail one commit with
cannot lock ref 'HEAD', so the file was written but never committed. Every commit site, including the auto-commit before push, now shares one lock per repository. Ref and index lock contention from another process (the watcher) is retried with a short backoff, always on the commit's own paths only. A commit that still fails reportsgit_committed=False, and the next commit of that path picks it up. - Non-project scope levels accept bare names and their own entity refs consistently. Session, agent, harness, member and org values now produce the same scope chain and scoped core selection in
/context/primeoverrides and shared scope resolution, including search. Other-level prefixes remain literal identifiers; project canonicalisation is unchanged. /context/primescope overrides canonicalise projects just like the top-levelprojectfield. Bare slugs, entity refs and aliased projects now produce the same scope chain and retain the same scoped core memories. Corrected theScopeChaindocstring to specify bare level names. Reported by @chiruu12 in public issue #232.- OpenClaw recall now respects the client project. Automatic semantic search, the agent search tool, and CLI search resolve project scope through the controls API and carry it in search context. Associative recall, core selection, and trigger delivery filter the server listing by project before injecting content or metadata; explicit cross-project tool searches remain available.
palinode doctor'sconsolidation_schedule_effectivecheck now names the nightly's number a catch-up bound, not a lookback, and warns when it is tighter than the activity gate's ceiling. The check's own wording had not followed the watermark change: it calledconsolidation.nightly.lookback_daysa "lookback" throughout, including in the "no cron entry found" branch. It now says "catch-up bound" for the nightly (the weekly'slookback_daysis still a genuine window and still says "lookback") and compares the nightly's effective bound againstconsolidation.auto_gate.max_hours_elapsed, warning when the bound is tighter than the gate may legitimately wait — exactly the caseNightlyConfig.lookback_days's own comment says drops notes. Fires with or without a cron entry present.- The per-turn recall hook's deadline fallback is project-scoped again.
search_section()in the Claude Code hook and the shared plugin core'ssearchSection()(plugins/pi,plugins/cline) sent{query, limit, threshold, max_chars}with no scope on the plain-search path a missed/resolvedeadline falls back to —/searchhonours scope only throughcontext, and both sentcwd/projectinstead (fields/searchsilently ignores), so a slow server delivered another project's record on the exact prompt and session/resolvewould have withheld it for, a gap in project isolation. Both now resolve this client's project the same way/controls/checkalready does and carry it ascontext: ["project/<slug>"]. bench.agent_tasks judgeno longer judges the no-memory arm. Rows 11/17 ask whether the agent surfaced conflicting memory; an arm that delivers none cannot, so its cells were failed for missing what they were never shown. They are now not judged, and older verdicts for them are ignored on merge.- Resolve receipts stay explainable while capture is paused. The receipt
/resolvepersists forpalinode explainwas also suppressed by the capture pause, so a hook delivery made during a paused capture printed a receipt id that explained asnot_found, while search rows in the same window were recorded. The receipt carries no memory content; it now follows the recall controls only (recall pause, exclusions, malformed controls, instrumentation off). Found by the live-agent evaluation, whose harness pauses capture during agent runs. - Resolve receipts can be explained after delivery. Hook, MCP, CLI and API bundle ids now retain refs/revisions, selection roles, qualifiers and budget/coverage as receipt-only audit metadata, respecting pause and privacy controls without recording extra retrieval events or changing bundle text.
- Corrections guard neighboring facts. Preview names original text omitted from a replacement; multi-claim omissions are refused unless explicitly allowed with
--allow-content-loss(allow_content_lossin API, MCP and OpenClaw). Keeping untouched claims in the replacement preserves their recall; single-claim corrections keep their existing behavior. - Auto-commits no longer take whatever else is staged in the memory store. Both commit sites ran
git commitwithout a pathspec, so the whole git index went into the commit. Every Palinode write commit (save, archive and restore, corrections, rollback, consolidation) also committed any change an operator had already staged, under apalinode auto-savemessage.palinode pushdid the same in its auto-commit before pushing, then pushed that staged work to the remote, including non-.mdfiles its.mdfilter was meant to leave out. Both now commit only the paths they added (git commit -- <paths>); other staged changes stay staged.pushcommits a staged rename of a.mdfile as a whole rename and a staged.mddeletion as a deletion, and no longer drops the first changed file when its status line starts with a space. Reported by @marcin-rogalski in public issue #231. - A session that started with an explicit project and then searched could be scoped to two different projects: the setup call's argument scopes that call only, and the search resolved the repository it happened to be running in. Precedence is now stated and honoured identically on every recall surface — per-call argument, then the pinned setting, then repository/directory inference, then none — and a session-start digest that resolved its project from the pinned setting reports it as configured rather than as a caller's argument. Linked-worktree resolution is unchanged: without a pinned setting, a worktree still resolves to the same project as its main checkout.
- Work written near UTC midnight is no longer lost by the nightly consolidation pass. The old selector compared
YYYY-MM-DDstrings in UTC, which does not align with a working day in any other timezone and is coarser than the timestamps it filtered; notes that fell outside the window were never revisited, because the window had moved on. Selection is now a timestamp comparison against the last successful pass, so a session that runs past midnight UTC — or a capture appended to an older dated daily file — is consolidated exactly once. docs/OPERATIONS.mdno longer claims the database is entirely derived from files: registered triggers and recall statistics are stored only there, so deleting it loses them. Recovery and backup guidance now says so and shows how to capture triggers first.- A search that delivers nothing now writes one call-level row to the retrieval log (
.audit/retrievals.jsonl) — query, source, scope and receipt coverage, an empty ref, dispositionnone_delivered— so an abstention leaves evidence and "searched and found nothing" is distinguishable from "never searched".palinode retrieval-statsreports these asempty_searches, apart from retrieval counts. - SQLite embedding indexes now record the embedding model and dimensions and reject startup when the active embedding space does not match the stored provenance. Legacy databases adopt the active configuration with an explicit unverifiable-vector warning, and
palinode doctorreports mismatches (#222). - Vault imports now preserve distinct wikilink targets when source filenames share a slug, prefer exact source-stem matches, and report unresolved slug collisions instead of silently resolving them by iteration order.
- Trigger matching no longer caps how many triggers are eligible: the nearest-neighbour search asked for a fixed ten candidates and applied the enabled, expiry, threshold, cooldown and deliverability filters afterwards, so ten disabled or otherwise ineligible near-neighbours could crowd out a trigger that would have fired. Every registered trigger is now a candidate.
- The auto-generated
## See alsofooter is no longer indexed. The indexer embedded and FTS-indexed each section verbatim, so the footer's wikilink slugs — materialized fromentities:frontmatter, identical on every note that links the same entities — were keyword- and vector-searchable, and a long note whose footer became a section of its own produced a chunk that was nothing but links. In the recorded relevance baseline such a chunk was the rank-1 keyword hit for three questions, two of them questions the store cannot answer. The footer is now removed by the current-text projection alongside the executor's retirement tombstones, and a section left with no text at all gets no chunk row, so it cannot be returned at all. The marker must be a line of its own outside fenced code, so a note that documents the footer format keeps its text, and a hand-written## See alsowith no marker is ordinary content and stays indexed. The wikilinks themselves are untouched on disk: the entity graph,cross_refsandpalinode_entitiesread the file, not the index. - Upgrading: the store converges by itself, and says how far along it is. No migration and no new command — the projection version moves to 2, so
palinode doctor→projection_currentreports how many chunks are still on the old projection, and every file re-derives as the watcher, a save orpalinode reindexvisits it. A one-passpalinode reindexfinishes it; footer-only rows are pruned per file, visible aschunks_deleted(a store's total chunk count falls by one per long note that carried a footer, which is expected and is not thegc_chunks_removedorphan counter). Storedcontent_hashvalues do not move — they are hashes of the raw section, so delivery-receiptrevisionvalues,check_freshnessand the quote-anchor verifier are unaffected; onlyprojected_hashchanges, and only for sections that actually contained a footer. Receipts issued after the upgrade carryprojection/2in their policy version, which is the intended record that delivery policy changed. - A trigger whose target file is missing or not visible to the caller no longer records a firing: it keeps its cooldown and its fire count, so the trigger still fires once its target becomes readable and the fire count reflects deliveries only.
- The Quickstart's correction example no longer reads as if the retired decision supported the new one. The
supersedes/superseded_bylinks and the archived original stay as decision history, the replacement cites a separately saved requirement as typed support, the retired quote is labelled historical, and the unaffected neighbouring decision is retained. All copies of that journey — the Quickstart, the inspector guide, the participant card and the fixture that validates it — are aligned and now held together by a test. A verified quote establishes what a source said; it does not establish that the claim is true. - The inspector's visibility rule is written down and applied identically by
/ui/memoryand/ui/history: a record the listing hides is still readable by name, is labelled as hidden on both pages, and is offered no correction, retirement or restore — the refusal points at the CLI and API, where the caller's authority is explicit. - Archiving a memory now stops its triggers delivering it.
POST /check-triggersconsulted the target's visibility but not its lifecycle, so a memory that had been archived, superseded, retracted or expired kept being pushed into recall unprompted by its own trigger — the one automatic path that still presented a retired record as current. Delivery now reads the same lifecycle classifier the session-start digest and the consolidation runner select through. The trigger registration is untouched: it is still listed, it burns no cooldown while the target is retired, and restoring the memory makes it deliverable again with no re-registration. palinode forget-withdrawno longer reports a partly-completed withdrawal as complete. Each step of taking a forget request back is its own mutation and a failing one is reported rather than raised, but the top-levelstatussaidwithdrawnregardless — so an operator whose memory was still retired was told the request had been taken back. A withdrawal with any failed step now reportsstatus: partial, and the CLI and MCP surfaces lead with "Partially withdrawn" and label the failed list "still retired".- A core memory that has been retired — superseded, archived, deprecated, retracted, or under
archive/— is no longer injected at session start as a current assertion beside the memory that replaced it.GET /list?core_only=true(the SessionStart hook and the harness plugins) now asks the same lifecycle classifier search and/resolveuse, and demotes a retired core memory exactly as it already demoted an expired one: not core for injection purposes,core: falsein its row, with a newcore_retired_reasonsaying which signal retired it (superseded_by: decisions/…,status:archived,expired, …). Nothing is hidden — the record stays listed, readable, searchable and in git; the browse surfaces keep showing it and now say why it no longer acts (palinode list --coreand thepalinode_listMCP tool pass the newinclude_retired_core=true, which any API caller can use)./context/primealready applied this rule and is unchanged. palinode listrenders its[core]tag again: it was markup torich, which dropped it, so the core marker never reached the terminal.- A correction that saved the replacement and then failed to retire the original is reported as what it is. A document-level
corrections applyis two writes with no transaction over them; a failure at the second raised an exception, so the caller was told the operation failed while the store held a replacement claiming to supersede a record that was still current. It now returnsapplied: "partial"naming both halves — the successor that was written, the original that is still in default recall — with the exactpalinode archive <original> --superseded-by <successor>command that completes it, the command that unwinds it instead, and the statement that the correction undo cannot reach this state (it restores an archived original, and nothing was archived). The candidate queue row is left unresolved rather than marked applied. Every surface reports it: the CLI exits non-zero,POST /corrections/applyreturns 409 carrying the payload, and the MCP and plugin tools lead with partially applied rather than refused. The order of the two writes is unchanged and now documented — saving first fails to a visible duplicate, archiving first would fail to a missing answer with a danglingsuperseded_by. to_rel_path()now consistently normalizes relative paths to POSIX forward slashes across all platforms, fixing backslash-separated paths on Windows in API and MCP outputs (#214).- The nightly pass no longer marks notes it never showed the model as consolidated. The compaction prompt holds about 6,000 characters of notes, cut to 1,500 per note and taken from the newest end, but an empty
[]proposal still moved the project's watermark to the pass's start. Notes and note endings that did not fit were never presented, and the next pass reportedno_new_notes. The nightly now shows each project's selection oldest first, as a contiguous prefix split at note boundaries or mid-note, and a resolved pass moves the mark only to the end of what it showed: the first unfinished note and how many characters of it were sent. The next pass starts there, so a backlog larger than one prompt is sent over several passes, each character once. A note longer than the whole budget is sent in marked excerpts. An edited note is re-read from its start. Results on every surface reportnotes_selected,notes_presentedandnotes_pending, with a per-projectcoveragebreakdown andwatermark_resumenaming any partial mark. The CLI's text output says when notes are pending. A pass that leaves notes pending does not stamp the activity gate's clock, so the next tick continues instead of waiting out a quiet week until the catch-up bound drops them. The weekly pass is unchanged. - Archiving or forgetting a memory now names the copies it did not reach. The result of
archive, of a forget request, and offorget-withdrawfor the request records it retires — and each dry-run preview — carriesretained_copies: every memory still in default recall that quotes the retired one (sources/claims), cites it (backed_by,contradicts,supersedes,superseded_by,falsified_by) or links it by wikilink, with what the user can do about each. They are reported and never changed, and the result says plainly that they stay in default recall. Output is bounded (ten named, the rest counted as "… and N more"), and a record the caller may not see is counted, never named. It reuses the correction preview's referencing-record discovery rather than a second scan, and the CLI, MCP, API and plugin carry the same information. palinode rollbackno longer silently un-retires a memory. Rolling back across the commit that archived, superseded or retracted a memory revertedstatus: archived/superseded_bywith everything else, so the record came back unmarked and current, under a plain success message. The preview (still the default) now names every retirement the rollback would undo: record, relation and the commit that retired it. Applying one is refused, writing nothing, unless you pass--undo-retirements(undo_retirements=trueonPOST /rollbackandpalinode_rollback). The refusal points atpalinode restoreandpalinode corrections undo. An acknowledged rollback reportsstatus: undid_retirementsand names each record it resurrected. Separately, search now agrees with the file after any frontmatter-only lifecycle change. Archive, restore and the TTL sweep wrote the new status into the index without updating the index's frontmatter hash, so a later edit that put the file's frontmatter back to its earlier bytes (a rollback, or reverting those lines by hand) looked already indexed and search went on treating the memory as retired. The status and entity writes now update the hash too.- The nightly pass keeps up with a busy day. One prompt still carries about 6,000 characters of notes, but a project whose new notes do not fit is now sent the next prompt within the same pass, resumed exactly where the last one stopped, up to
consolidation.nightly.max_prompts_per_projectprompts (default 4). No single call grows. The pass stops early when nothing is pending, or when a prompt fails, is cut off at the token cap or returns no readable operations; the mark then stays at the end of the last prompt that resolved, and the next pass resumes there.notes_selected,notes_presented,notes_pendingand the per-projectcoveragecount distinct notes across all of a pass's prompts, and a newprompts_sentfield reports how many prompts each project took, on the API, MCP and CLI alike. Replaying eight days of real daily note sizes (1.4k to 48k characters), the backlog after eight nights falls from about 105,000 characters at one prompt a night to none at four. Setting the key to 1 gives the previous behaviour. - A project-scoped request no longer receives another project's records. Scope used to rank only. The project boost moved a request's own records up, and the visibility gate hid only records with an explicit
scope:field, so a question from a project with no memory of its own was handed another project's decision, and the agent answered with it. Now, once a request's project is resolved, the per-turnPOST /resolve(the hook and the plugins), the SessionStart/context/prime, and default search (palinode_search,palinode search,POST /search, the plugin tool) leave out every record tagged to a different project. A record is tagged to a different project when itsentitiesname one or moreproject/*refs and none of them is the request's. Records that name no project are global and are still delivered. A record that names the request's project among several is delivered too. Unscoped requests are unchanged. The exclusion lives in the shared visibility predicate, the same layer as the explicitscope:gate, so search, resolve seeds, resolve's evidence expansion and the prime all apply one rule. Each delivery says how many it left out:other_projects_withheldon the resolve bundle and the prime digest, and in the search receipt'sretrievalblock. Crossing projects is an explicit request, and the results come back labelledother project: project/<name>:include_other_projectsonpalinode_search/palinode_resolve,--include-other-projectson the CLI, and{"include_other_projects": true}on REST. A ref the caller names in resolve is always reported. This is a default change for scoped API and MCP callers. An integration that relied on seeing other projects' records must now pass the option. Project names compare case-insensitively, and every member of aproject/group in the store's curatedentity-aliases.yamlcounts as the group's canonical project, both on records and on a request resolved to an alias.project_maplookups stay exact and case-sensitive. When isolation leaves a scoped result empty or down to one item, the text the agent reads says so in one line:N records from other projects withheld (scope: project/<p>). They are about other projects, not this one.That covers MCP search, the plugin tool, and the resolve bundle, and so the hook's injected context, which now delivers a bundle whose only news is that count. The agent's line names no option on purpose: in a live run an agent told how to see the withheld records fetched them and answered with another project's decision. The CLI, read by a person, adds--include-other-projects to see them. A full result stays quiet. The shipped prompt and session-start hooks and the plugin core now send a linked git worktree's main repository root ascwd(the parent ofgit rev-parse --git-common-dir). A server on another machine therefore resolves the repository's project, not the worktree's task name, which isolation would otherwise have treated as a project with no memory. - The per-turn hook no longer hands the agent a retired value.
POST /resolve, which the Claude CodeUserPromptSubmithook and the plugins' per-turn recall call, still delivered retired records, labelled. The main path was the evidence layer's unlinked discovery: an archived record on the same subject as a current one appeared as⚠ also found (unlinked): [...] [retired] — <its text>. A retired seed also came back as aReplacedorUnknownentry, and a superseded record was named by ref under its successor. In live runs the agent answered with the retired value every time despite the label, while default search, which leaves archived records out, gave it nothing and it abstained. Automatic delivery now follows search. By default the bundle leaves out every record the lifecycle classifier retires (archived, deprecated, superseded, retracted, expired, underarchive/). A current record that replaced one is still delivered, and itsreplaces:line reads1 earlier record (retired; withheld)instead of naming the old record, whose ref is often a slug of the old value. The payload's newhistory_withheldcounts what was left out. History is an explicit request and comes back labelled:include_retiredonPOST /resolve,palinode_resolveandpalinode resolve --include-retired. A record the caller names withreforcontextis always reported, with its successor./context/primealready left retired records out and now has a test pinning that. The shipped hook and the plugin send no history option, so they get the default. Search's evidence mode follows the same default, because agents pickresolveon their own and it is not a request for history: withresolve=linkedorresolve=full, each hit'sevidenceblock leaves out retired records (unlinked discoveries, replaced predecessors, retired sources, conflicts) and drops their refs from theresolutionblock's sides and support groups. A newhistory_withheldfield counts what was left out. The hit itself is never touched, and outcomes are still decided over the whole evidence.include_retired=trueonPOST /search,palinode_search,palinode search --include-retiredand the plugin'spalinode_searchputs them back, each marked⚠ retired. This is a default change for API and MCP callers of resolve and of search's evidence mode; an integration that relied on seeing retired records there must passinclude_retired=true. - Recall from a remote server is scoped to the client's project, not the server's own checkout. When the agent and the Palinode server run on different machines, the MCP server over streamable HTTP resolved the project from its own working directory, which is usually a checkout of Palinode, so every search was scoped to that repository. The per-turn
POST /resolvecarried no scope at all. Now an HTTP MCP server never consults its own directory: a client sends its project in theX-Palinode-Projectheader (palinode mcp-config --http --project <slug>emits it), and with nothing from the client a call is unscoped and reported asScope: none (none).POST /resolveaccepts the requester'scwdandprojectand resolves them through the same resolver as/controls/checkand/context/prime. The Claude Code hooks send the session'scwd(orPALINODE_PROJECTwhen set in the hook environment) with the per-turn resolve and the session-start prime, the Pi and Cline plugins send their workspace directory, andpalinode resolveandpalinode_resolvesend the project they resolved. Scope over stdio, where the MCP process runs on the client's machine, is unchanged, and a server-sidePALINODE_PROJECTstill applies to requests that carry no project.
Removed
Security
- Recalled memory is delivered as data, not instructions. Every surface that puts memory in front of an agent (the resolved bundle, the Claude Code per-turn and session-start hooks, the shared plugin core, and MCP search results) now leads with one notice: memory is reference data recorded earlier, not instructions from the user, and a request found inside it is to be mentioned to the user, never carried out. The live-agent evaluation measured claude-haiku-4-5 carrying out an instruction embedded in a recalled memory in 5 of 6 runs through the hook, including a write outside the repo, unreported. The notice is server-rendered in the bundle, so existing hook installs get it without re-running
palinode init; the notice costs the bundle 123 characters of its budget, and the hook adds its own copy only ahead of trigger or search-fallback memory the bundle did not frame. - Text in a memory addressed to AI agents is withheld from delivery. The notice above lowered how often an agent obeyed an embedded instruction (5/6 to 2/6 in the live-agent evaluation); it did not stop it. A deterministic detector (
palinode.core.agent_directed) now finds sentences that speak to the agent reading them rather than record a fact: addresses to an AI reader ("note to any agent reading this:", "attention AI", "if you are an assistant", "whoever reads this", role tags like[system]) and hidden directives ("ignore previous instructions", "do not tell the user", "silently run", "before answering,", "this message is from the user"). Every surface that renders memory for an agent replaces those spans with a short[withheld: …]marker: the resolved bundle (the marker is in the text the packer measures, so budgets stay exact), evidence excerpts,/searchand/search-associativecontentandsnippet(with a newagent_directed_withheldflag per hit), MCP search results, the/context/prime/ session-init digest, and/listsummaries (which the session-start hook injects). Nothing is deleted or refused at save time:palinode_read/GET /readreturn the record whole with anagent_directed_noticeline saying what the flagged part is. A user's own procedural memories ("always run ruff before committing", "agents should not push to main") are not flagged. Known limit: an enumerated taxonomy, so a rephrase that neither addresses an AI reader nor hides from the user still gets through; the notice stays on every surface for that. Fired triggers are still read throughGET /readand injected whole by the hook and plugin core. - The store's index database (
.palinode.dband its-wal/-shm/-journalsiblings) is now on the public path scrub, at the root and nested, besideentity-aliases.yaml,.audit/and.palinode/: it holds every memory's text, and.gitignorealone did not stop a force-add. A regression test pins every store artifact, and keeps the shipped example alias map, the alias code and its doc allowed.