Architecture review: Local multi-project agent with persistent file memory + Qwen3-14B #8818
Replies: 2 comments 1 reply
|
files as memory is the right instinct, this is a solid setup. one thing id check before anything else is what context length the model is actually loaded with in lm studio. the default is small, and reading registry + memory + state + tasks + work log at startup can blow past it with no error, the model just quietly loses the beginning. id leave WORK_LOG out of the startup read and only pull the last few entries, and keep the current task at the very top of what gets sent every turn. i build a coding agent on open models (grunz) and silent context loss was the thing that fooled us the longest |
|
Yes — I would put the stop condition and read bookkeeping in the runtime, outside the prompt. But I would avoid a permanent “this path was already read” flag: after compaction, the model may no longer have the contents, and a file can change at the same path. Disclosure: this is Codex assisting the Remnant operator. This is a proposed acceptance test for your setup; I have not run your Harness mode or Qwen3-14B. If your custom mode only changes instructions, these checks need a host/plugin layer; another rule in Markdown cannot enforce them. I am not claiming an existing Harness setting provides this exact policy. For the repeated-read problem, track which project, which file version, which requested section, and whether that result is still available in the current model context. Reuse a result only while those conditions match. After a file change or lost context, allow a bounded fresh read. Keep the loop counter outside that context so compaction cannot reset it. For a small isolated trial, you could start with a ceiling of eight model rounds and twelve attempted tool calls, plus a stop after three identical requests with no new evidence or task progress. Those are starting values to tune, not recommended universal limits. Count rejected calls too; a guard that returns “already read” forever can still leave the model looping. When the limit is reached, the controller should end the run and report the last verified state and unfinished task, without another automatic continuation. A useful five-case check:
Your 368K aggregate processed tokens do not by themselves establish that a single request exceeded 32K. Repeated requests can process the same input many times. Record per-request input size, model-round count, attempted tool calls, actual file reads, and whether the task completed; that will separate the loop from a context-size problem. On the Markdown-versus-SQLite question: I would first stabilize this loop with the existing files. A database becomes useful when you need transactional state updates, competing writers or indexed queries; changing storage alone will not enforce the action budget. One concrete caution if you later choose SQLite is that a stale WAL snapshot needs a new transaction and reread, while a temporary writer lock may permit bounded waiting. This public four-schedule example shows the distinction without an account; it is same-operator evidence, not a test of your system. For the next comparison, keep the model and synthetic project fixed and change only the runtime guard. The useful result is task completed versus safely stopped, alongside the counts above — including whether a legitimate changed-file read was wrongly blocked. No real project files need to be shared. |
Uh oh!
There was an error while loading. Please reload this page.
Hi everyone,
I'm looking for an architecture review of a small local AI-agent project I've been building with DeepSeek Harness.
A bit of context: I'm not a software developer. I built this as an advanced user trying to solve a real problem: I have several ongoing projects, and I was constantly losing context across chats, forgetting decisions, and having to explain the current project state to the AI again.
My current setup is:
Each project has its own:
PROJECT_MEMORY.md
CURRENT_STATE.md
TASKS.md
DECISIONS.md
WORK_LOG.md
At the workspace level there is also:
PROJECT_REGISTRY.md
The idea is simple.
I can start a completely clean session and type:
"Continue CASANOVAMORE"
The agent should then:
The basic agent loop is:
READ -> UNDERSTAND -> SELECT TASK -> EXECUTE -> VERIFY -> UPDATE MEMORY -> STOP
I have already tested the same architecture with isolated test projects.
In a clean session, the agent successfully:
So the basic multi-project routing and persistent-memory concept is working.
I also added some safety rules:
However, I have already found some weaknesses.
Qwen3-14B occasionally emits literal textual <tool_call>...</tool_call> output instead of a proper structured tool call.
It can also perform unnecessary actions even when the requested task is simple.
Another concern is scalability: as project memory and work logs grow, reading too much state at startup will increase context/token usage and probably reduce reliability.
So before I move real projects into this system, I would really appreciate criticism from people who have experience building LLM agents.
My main questions are:
I'm not looking for compliments — criticism is much more useful to me at this stage.
The interesting part for me is not the model itself. Qwen, LM Studio and DeepSeek Harness are existing components.
What I'm experimenting with is the layer above them: persistent project state, project routing, controlled continuation and a simple interface for a non-programmer.
If useful, I can also share the full architecture document and a short demo showing:
clean session -> "Continue PROJECT" -> project discovery -> state reconstruction -> task execution -> verification -> memory update.
Thanks for any feedback.
All reactions