EYAS v0.8.16-beta — A memory of its own
What EYAS knows now comes from what EYAS remembers. The vault writes itself: a
durable fact stated in any conversation — on any model — becomes a note without
anyone asking, and the same note is what every later conversation reads back.
Closing that loop meant winning an argument with the host machine: conversations
on the Claude Code CLI no longer read the owner’s own Claude config and memory,
because an assistant that can see a second memory will happily report a fact
“already recorded” that its own vault has never held.
Every fix here was found the same way: a live test, a measurement table, and a
root cause chased until it reproduced deterministically. The capture run ledger
(memory_capture_runs) is why each diagnosis took minutes instead of days.
Memory that fills itself — and knows where it came from
- A durable fact learned in a conversation is written to the vault without
anyone asking. Capture runs on every conversation, globally, on by default;
memory.capture.enabledinconfig/default.yamlswitches it off. A small
model call attaches to a qualifying turn AFTER the reply has been delivered —
never in its critical path — and a capture that fails is a missing note, never
a failed conversation. - The extractor reaches a model that can answer it, whatever the instance
runs. Capture assumes nothing about what is installed — most instances are a
VPS or a pod with no room for a local model, and many have nothing but a host
CLI. Resolution is a ladder over what is actually enabled: theheartbeat
tier, but only when this instance really has the provider that tier names and
it is not a CLI; otherwise the first enabled, registered provider that is not a
host CLI whose model can be named; otherwise no pin at all, letting the gateway
fall back the way it does for any unpinned request — anthropic when registered,
else the first registered provider, a CLI included — because a capture that is
attempted is measured and one that is skipped is invisible. The rung is
logged. The routing tier is configuration, so it can name a provider this box
does not have — and it did: the pin was silently dropped, and a CLI provider's
complete(), which runs a full agent turn, answered the extraction prompt in
prose. Oneunparsablerow per qualifying turn, and never a note.
Three repairs meet the CLI there as well: the parser lifts the first balanced
{…}object out of surrounding chatter (string-aware, so a brace inside a
value does not close it), the prompt says the reply is the object and nothing
else — no commentary, no fence, no tool calls — and the unusable-output
warning now carries the reply's length and its first 200 characters, so the
next diagnosis is not blind. - The extraction runs in an isolated context, so a CLI's own loaded memory
cannot pre-empt EYAS's. A request can now ask to beisolated— no
filesystem settings, no CLI-native memory or config, no bridged tools, a
single turn — and Claude Code honours it whatever itsloadClaudeMdsetting
says. Without it the extraction call loaded the owner's~/.claudememory,
which another tool had already written the fact into: the model read it there,
reported it known, and EYAS's vault — the one place it was NOT recorded —
stayed empty. No prompt rule wins against a whole loaded memory system. The
ladder now prefers a CLI that advertises the capability over one that does
not, choosing on the CAPABILITY and never on a provider name; grok CLI's
protocol offers no such switch, so it says so rather than pretending. - The extractor believes the notes on file, not the assistant's account of
them. A retest caught it returning a healthy-empty batch on a fact-dense
exchange: the reply had said "I've already saved that to memory" — it had not,
the CLI narrated a tool call that never ran — and the extractor honoured the
do-not-restate rule against that claim while its own EXISTING NOTES section
was empty. Coverage is now judged ONLY against EXISTING NOTES, and the prompt
says in as many words that an assistant's statement about saving is narration,
not evidence. What a model concludes cannot be asserted in a test; that the
instruction ships is pinned by one. - Every capture run records which model produced it.
memory_capture_runs
gains aprovidercolumn holdingprovider/model— NULL when no model was
called, because a gate skip spends nothing. An instance with several providers
could already count its unparsable runs but could not say which model was
failing to answer in JSON, and answering that took a live retest once already. - Memory is EYAS's own, in both directions. The mandatory memory rule named
no tool and only one direction ("update memory when you learn something new"),
which a CLI-backed agent reads as its own machine-global convention. It now
namessearch_memoryfor recall andsave_memoryfor recording, states that
EYAS's memory is the only memory, and forbids writing to a machine-global
memory directory, anai-memoryor Obsidian vault,~/.claudeor~/.grok.
Because a rule is guidance, the deterministic gate denies the same paths to
every file-writing tool, matching the path fields of a call and never its
content.Read,GrepandGlobstay open — the data-port importer exists
to carry exactly those notes into EYAS — but the shell is blocked in both
directions, becausecatis one character from>>and no reading of a
command string proves which one it is. The denied set is narrow on purpose:
anai-memorydirectory, a home-anchored~/.claudeor~/.grok, and a
memory/directory under either. A workspace's own.claude/settings.json
and.claude/agents/*pass, since that is project config, not memory.
MEMORY.mdis deliberately not on the list: the gate is handed a path, not a
workspace root, and cannot tell the owner's global index from a repository's
owndocs/MEMORY.md. - The gate is structural, not lexical. One length check,
minUserChars
(default 40), counted in Unicode code points so an accented message gates
identically to an ASCII one of the same length. No keyword list in any
language: this product ships in six, and the repository has already paid twice
for that class of bug — JavaScript's\bis ASCII-only, so\bűrlapnever
matched "Űrlapelemek", and Hungarian lengthens the stem vowel, so "minta" is
not a prefix of "minták". Deciding what a sentence MEANS is the model's half
of the design. - The runaway guard counts model spend, not turns.
maxPerConversation
(default 20) is consumed by a successful extraction, an unparsable reply and
an errored call — never by a too-short skip. Counting skips meant twenty short
acknowledgements ("ok", "mehet") exhausted the budget without a single model
call, and the next fact-rich turn was refused. Every outcome still writes its
row; only what the budget is spent on changed. - 0–2 candidate notes against a strict schema.
user(who the owner is),
feedback(how to work — invalid unless it carries both a Why and a How to
apply),project(a durable fact about the conversation's project) and
reference. When the conversation has no real project, aprojectcandidate
is REJECTED by the schema rather than hidden from the model — and because the
refinement runs per note inside one array parse, a single stray project
candidate fails the whole batch, which is then dropped and recorded as
unparsable.{"notes":[]}is the common and correct answer, and the prompt
says so. - A repeated fact reinforces one note instead of spawning a second.
Deduplication is word-set overlap against the existing summary rather than
string equality, because a reinforcement rephrases ("Answers in Hungarian" →
"Answers in Hungarian, always"); a match appends a dated bullet under
## Historyand never overwrites what was there. - Sanitised before it touches disk, not when it is read. The privacy module
runs over the summary and the body before the vault write, because a read-time
redaction would leave the raw text in the file and in the FTS index built from
it. - A project's facts rank first inside that project and are invisible outside
it. The always-on index ranks globaluserandfeedbackfirst, then the
ACTIVE project'sprojectnotes, thenreference; another project's notes
never appear at all. Project notes live inprojects/<project-id>/with a
projectfrontmatter field frozen at capture, so re-scoping a note is a
deliberate act rather than a side effect of the next update. - The seed catch-all project is not a project identity. Every conversation
defaults intogeneral-general, so treating it as a real project would file
the owner's general facts under it and hide them everywhere else.
The rule lives in one FUNCTION,effectiveProjectId(), and every entry point
calls it — capture, both recall paths, and the memory tools — so the write
half and the read half cannot disagree about what counts as a project. - Every note records where it came from.
memory_note_linksnames the
conversation that wrote a note or later reinforced it, in the same multi-owner
shape asdesign_linksanddocument_links, and episodic memories now carry
conversation_idandproject_id. - Every outcome that reached the gate writes a
memory_capture_runsrow —
skips with their reason, extractions with the kinds they wrote. Two silences
are deliberate: capture switched off writes nothing at all, and a background
run with no assistant text to read never reaches the gate, because a skip row
per autonomous run would only inflate the diagnostics it exists to keep
honest. One silence is a known gap rather than a choice: a God Mode turn
returns its own stream before the post-turn block and so captures nothing —
no note, no row. That is what makes
minUserCharsa tunable number instead of a permanent guess, and thekinds
distribution is the only way a mislabel (a project fact filed asuser)
becomes visible without reading the vault by hand. - A boot-order bug had silently stopped the memory lifecycle hooks from ever
wiring.conversationscheckedctx.memoryin its ownonStart, but the
loader orders modules by harddependenciesonly and startsconversations
beforememoryon every boot — so the hooks were wired 0 times in 68 recorded
start cycles. They resolve lazily now, with the same pattern as the lazy
getters that were already on the lines above the old site, and PreCompact
summaries reach episodic memory.
Claude config isolation is now the default
loadClaudeMddefaults to OFF. Conversations on the Claude Code CLI no
longer load the host machine's Claude config — nosettings.json(hooks,
permission rules), no CLAUDE.md at any tier, no host skills, no project
.mcp.jsonservers.
(Enterprise-managed policy settings, where deployed, still apply — that tier
cannot be suppressed client-side.)
EYAS's own memory is the single source of truth; what the model knows is
what EYAS recorded. Existing installs flip too — one click on the provider
panel opts back in.- The ON path is explicit now. Opting in sends
settingSources: ['user','project','local']instead of omitting the option.
The installed CLI treats an absent flag as "load everything" while the SDK
docs promise the opposite — the toggle no longer depends on either reading,
and the panel copy says honestly that ON loads the whole machine config,
hooks included, not just CLAUDE.md. - No fake switch for Grok/Kimi. ACP has no isolation parameter and the
grok CLI has no suppression flag (it demonstrably loads~/.grokand even
~/.claudeglobally); the kimi baseline is unverified. Their panels now say
so instead of pretending otherwise. - Known residual: a CLI session created before the flip restores its
previously loaded context when resumed, until the session goes stale. - Auto-memory and filesystem MCP configs are covered too.
settingSources: []
alone was not the whole story: the CLI's auto-memory keys the machine-level
~/.claude/projects/<cwd>/memory/MEMORY.mdon the working directory — a live
extraction read the owner's own memory index there and judged a fresh fact
"already recorded". Isolated calls and opted-out conversations now also set
CLAUDE_CODE_DISABLE_AUTO_MEMORYandstrictMcpConfig, closing both channels.
Known issues
- The capture gate's threshold is still a guess — but now a measured one.
minUserChars: 40was chosen before there was any data on how often it fires;
memory_capture_runswas built so the next change to it is read off the skip
distribution rather than argued. The same table'skindscolumn is the only
signal on whether the extractor mislabels a project fact as a fact about the
owner, and neither has enough rows yet to say.