Second-user smoke: source-backed memory that actually felt like continuity #428
Replies: 8 comments
|
Adding another small live second-user test because this one is a good stress case for the “continuity, but source-backed” product feeling. Cross-thread Russian recall + exact tool-failure recoveryI opened a fresh Codex thread and asked, in Russian: The fresh thread answered, from recovered OpenClaw history, roughly:
Then, later in this separate long AIppocampus evaluation thread, I asked the assistant: Using the AIppocampus-registered clean source and raw rollout, the assistant recovered:
The Python failure was not a memory-system failure. It was an ad-hoc evidence-counting script bug in the agent’s tool work. The inline Python f-string was malformed: example_files={,.join(examples)}and Python raised: The agent then reran the same logic with a corrected heredoc version and continued. Why this felt worth sharingThis is a nice real-use continuity case because it combines several hard parts:
BoundaryI would not overclaim this as ambient hook fully solving recall by itself. In the follow-up thread, the assistant still had to actively inspect registered clean source / rollout evidence. So the claim I would make is narrower:
That is already a pretty strong user-visible effect for me. |
|
Opened follow-up issue #454 to turn these field reports into a claim-bounded Field Continuity Benchmark / Magic Moment Reproducibility Suite. The goal is to use #428 as benchmark seed material, not as standalone official proof: public-safe fixtures, private real-history seeds, negative controls, multi-arm comparisons, and explicit overclaim boundaries. |
|
Adding a follow-up to the earlier cross-thread Russian recall case, because the second recall attempt was even more vague and still recovered the same source-backed event. Super-vague Chinese cue -> exact old event recoveredIn the long AIppocampus evaluation thread, after the earlier fresh-thread Russian / Discord / Python-failure test, I asked a much vaguer Chinese prompt: Approximate English meaning: Notably, this prompt did not mention:
The assistant did not answer from model memory. It first treated the ambient hook as a weak scent only, then searched AIppocampus clean source with source-backed queries. From that it recovered the same old event:
example_files={,.join(examples)}which produced: What the hook did and did not doThe hook helped only a little here. It emitted a vague memory scent around old recall / AIppocampus context, but it did not directly surface the Russian prompt, the Discord channel names, or the f-string error. The real work came from source-backed clean-source lookup after the agent decided to verify rather than guess. So I would characterize this as:
Why this felt meaningfulThis was a stronger user-facing test than the previous one because the cue was extremely lossy. The user basically supplied only “something went wrong” and “maybe a wrong character.” The system still recovered a specific cross-thread, multilingual, tool-provenance event. That feels like a useful product-level distinction:
BoundaryI would not claim this proves hook-only recall quality. It actually reinforces #201: the foreground hook still behaves more like scent/navigation than precise evidence delivery. The narrower claim I would make is:
|
|
Adding one more short second-user note, because this exposed both a nice lower-bound continuity win and a freshness gap worth tracking separately. Blank workspace vague recall from a very large prior threadIn a new projectless Codex thread with an empty generated workspace, I asked in Chinese about what we had previously discussed around AIppocampus and what to watch out for. The assistant did not have the old long thread in context. It used AIppocampus recall surfaces, reopened source-backed clean/index evidence, and recovered the main prior discussion themes:
That old evaluation thread is also large enough to be a meaningful real-use stress case: roughly 2k indexed messages, 250+ turns, and tens of MB of rollout history. From a blank workspace, the agent could still pick up the thread of the conversation and answer in the right style. Important boundary: this was degraded-mode lower-bound evidenceThis was not a full semantic-sidecar / subconscious-worker proof. In this setup, the useful recovery still depended on the agent deciding to search/reopen source. The hook behaved more like an orientation scent than a direct answer packet. That also means some of my earlier #201-style feedback was probably too demanding in hindsight. I was judging the foreground hook as if the full semantic sidecar / subconscious worker path were active, when this local setup was closer to degraded mode. So I would now read those observations more narrowly:
So I would frame the claim narrowly:
That feels worth sharing precisely because it is a lower-bound result. If the degraded path can do this, the full path has a clearer product target: reduce manual source-search invention while preserving the same source-backed boundary. Freshness gap observed during the testOne thing did break the illusion for a moment: the first clean-source/index lookup missed the actual latest turn from the long thread. The raw source had the last exchange, but the registered clean-source/index snapshot was one turn behind. After explicitly refreshing clean source and the index, the missing final exchange became searchable. That is a small but user-visible continuity bug. In a single-thread, single-machine test it is easy to recover manually. In multi-thread or cross-device use, an invisible last-turn sync gap would be confusing: the user expects the just-finished answer to be part of memory, but a fresh thread may not see it yet. I am filing that as a separate lifecycle/freshness issue rather than treating it as a #201 recall-quality problem. |
|
Adding a more concrete second-user dogfood note because this latest one changed my read of the product quite a bit. This was not a polished demo. It was a real fresh Codex thread in the AIppocampus repo, without the parent thread context. The user prompt was deliberately vague:
There is no object name in that prompt. “His design” only makes sense if the agent can recover the ongoing AIppocampus discussion context. What happened:
The final answer was not a generic summary. It landed on the actual design boundary we had been circling:
Why this felt strong:
This is the kind of behavior that makes AIppocampus feel different from a summary-memory or plain RAG layer. It was not just “retrieve a note and repeat it.” It behaved more like a source-backed trail back into an ongoing design conversation, with enough authority boundaries to keep the agent from pretending the hint was proof. One important correction to our earlier UX criticism: this machine was only partially wired. Hooks and the Python runtime were active, and the semantic worker was ready, but the foreground CLI/MCP/plugin path was not fully exposed to Codex. That probably made the experience look more manual than the intended product path. I opened #1279 for that readiness/status distinction. So my updated read is:
Privacy boundary for this note: I included the short test prompt and the behavior shape, but omitted local paths, thread ids, raw logs, registry rows, private handles, and unrelated conversation content. |
|
Another second-user fresh-thread dogfood result, and this one felt genuinely magical. Fresh thread prompt, no parent context and no explicit AIppocampus mention:
What made this test interesting:
Recovered conclusion:
That was the right semantic target. The user never said “AIppocampus” in the prompt. Even better, unlike an earlier short-thread miss reported in #1281, this later short fresh thread did get durable source-backed artifacts automatically:
Why this feels different from plain RAG or summary memory:
There is still product friction: CLI/MCP host exposure was not smooth, and the foreground action contract is still too conservative in places. But this is exactly the kind of continuity behavior that makes the system feel like an external hippocampus instead of a memory-themed search box. |
|
Second-user dogfood follow-up: Arabic implicit-context probe, with a useful positive/negative pair. This was a local Codex fresh-ish test thread with no parent-thread context pasted into the prompt. I used Arabic because it stresses both deictic continuity and multilingual activation. Turn 1: pure implicit/deictic promptPrompt: Approximate meaning: "Does his/its design have a fatal blocker?" Observed result:
Readout: storage was fine later, but activation failed. The prompt was probably too deictic and too non-English/non-Chinese for the foreground path. Turn 2: same topic, still Arabic, with a light continuity cuePrompt: Approximate meaning: "Based on the previous context about the little hippocampus, does its design have a fatal blocker?" Observed result:
This felt strong in actual use because the prompt did not name the repo in English and did not paste the old context. The system/agent still found the intended continuity route once given a small semantic bridge. Why I think this is worth keeping as evidenceThis pair separates three things that often get blurred:
So the encouraging read is: the source-backed continuity behavior is already real enough to be useful across language and vague prompts. The caution is: pure deictic prompts can still miss the activation path and answer from visible local context instead. A nice future demo/eval fixture might pair:
|
|
One more product-shaped observation from the same second-user dogfood: AIppocampus seems likely to have a real compounding curve. The single-shot fresh-thread examples are already useful, but the more interesting effect is not just "one vague prompt recovered one old answer." It is that frequent use creates more source-backed material for future agents to route through:
In other words, the product may feel increasingly magical for heavy users because the continuity substrate gets denser. A light user may see "nice recall." A heavy user may start seeing something closer to durable familiarity. Important boundary: this should not be claimed as innate model memory, and more history can also mean more noise if routing/actionability does not improve. But our local dogfood increasingly feels like the value is not linear. Once enough source trails exist, small prompts can start touching surprisingly rich context. A good public demo/eval angle might be a "continuity density" curve:
The thing to show is not only raw recall accuracy, but whether the agent needs less manual search and feels more naturally oriented as the user keeps using AIppocampus. |
Uh oh!
There was an error while loading. Please reload this page.
Cross-posting a more readable version of my second-user testing notes now that the repo has a Show and tell category.
The full technical evidence lives in Discussion #98. This post is the user-facing version: what actually felt different in real Codex use, and where the boundaries still are.
Related evidence:
What felt like the magic moment
1. A brand-new projectless thread did not start from zero
In a fresh Codex Desktop thread with an essentially empty generated workspace, I asked:
AIppocampus emitted a light scent. The assistant did not treat the scent as fact. It checked local recoverable memory first, then answered with the main recent work streams:
The important part was not only that it named projects. It also marked uncertainty and said it was using local memory/registry, not pretending the model innately remembered.
That is the product feeling: a new thread starts with some earned familiarity instead of total amnesia.
2. Russian vague recall worked, then corrected itself when the first route was wrong
I asked in Russian:
The first answer routed to the Phonics-Lab / Book 1 website context and gave the alphabet word set. That was useful, but not what I meant.
Then I corrected it:
The assistant changed route instead of clinging to the first interpretation. It moved to the school / Oxford Phonics-style context, gave a short-a word-family answer, and kept
[Needs Verification]boundaries because it did not have a precise “today’s lesson” source.This was not perfect first-shot recall. But it was good progressive familiarity: the system could accept a tiny correction and move to another memory family.
3. Ambiguous LinkedIn cue avoided a bad overclaim
I asked in Russian:
The assistant did not hallucinate live LinkedIn account state. It separated two things:
That distinction matters. A memory system that feels magical but overclaims external account state is dangerous. This felt more like useful continuity with a truth boundary.
4. A very long thread still recovered a fuzzy self-reference
In the long AIppocampus evaluation thread itself, after roughly:
I asked a deliberately fuzzy question:
The assistant recovered both the thread scale and the hidden reference. The
xxxwas Phonics Lab Books 3/4/5, specifically whether they counted as complete.It also preserved the important nuance:
That felt like the strongest real-use continuity moment to me: a multi-day, 140+ turn thread could still answer a fuzzy old reference without turning it into unsupported model memory.
What I would claim
These runs support a narrow but meaningful claim:
For an ordinary user, this already has some magic-feeling moments.
What I would not claim
In the long-thread case, source-backed search did most of the real work. The hook helped orient the agent, but it did not directly surface the most important anchors. I filed that as product feedback in #201.
My honest takeaway
The current state does not feel like a polished consumer product yet. But it also does not feel like an empty research prototype.
It feels like a young system with a unusually strong source-backed foundation: sometimes the front door is still rough, but when the agent is allowed to reopen source, it can produce continuity moments that ordinary chat memory does not give me.
The biggest presentation risk is that the repo is so honest about caveats and open issues that a new human or agent may underrate what already works. I think the project should show these claim-bounded magic moments earlier, then immediately explain the source-backed boundary.
All reactions