-
Notifications
You must be signed in to change notification settings - Fork 0
Architecture
CLI (REPL) Gateway MCP server Cron
cli.py gateway/ mcp_server.py schedule_task
│ │ │ │
└─────────────────┴──── run_agent() ───┴──────────────┘
engine.py
│
┌─────────────┼─────────────┐
check_permission run_tool MCP client
(the gate) (dispatch) mcp.py
The single most important structural fact: every front end calls the same
run_agent() and the same run_tool(). Adapters translate transport details
only. They do not re-implement permissions.
That's why fixing the permission layer once fixes it for the terminal, Slack, Discord, WhatsApp and MCP simultaneously — and why a front end that does diverge (as the MCP server briefly did by forcing full-auto) is a bug rather than a design choice.
| File | Responsibility |
|---|---|
engine.py |
Agent loop, tool dispatch, permission layer, security floors, providers, HTTP/SSRF |
cli.py |
REPL, slash commands, setup wizards, rendering |
providers.py |
The 12 built-in provider profiles |
mcp.py |
MCP client — connect external servers |
mcp_server.py |
MCP server — expose our tools outward |
gateway/runner.py |
Inbound routing, approval registry, adapter registration |
gateway/agent_bridge.py |
Bridges a chat turn to run_agent()
|
gateway/session.py |
Per-chat JSON session files |
gateway/auth.py |
Allowlist, per-platform scoping, WhatsApp id normalisation |
gateway/platforms/*.py |
Slack / WhatsApp / Discord / Telegram / Email adapters |
memory/store.py |
Memory: SQLite schema, hybrid BM25 + vector search, RRF |
memory/embed.py |
Memory: embeddings via the provider's /embeddings
|
memory/extract.py |
Memory: the fact-extraction prompt and its parsing |
agents/*.md |
Built-in sub-agent profiles (auditor, coder, explore, general-purpose, researcher, test-writer) |
open_code_review.py, review_store.py
|
review_code tool / /review: native + delegated review, stored findings |
repomap.py |
repo_map tool — dependency-ranked repository outline |
repository_access.py |
repository_read tool — local dir or github.com repo survey |
document_access.py, documents.py
|
document_read cached extraction; document write skill support |
local_models.py |
Ollama hardware probe / list / pull / remove tools |
web_search.py, web_server.py
|
web_search backend routing; --web browser-UI mode (FastAPI/uvicorn) |
logging_setup.py |
Daily-rotating JSONL operational log, attached to the agent8088 logger tree |
testing_support.py |
Shared test/verification helpers |
tools.txt |
Tool registry (data, not code) |
system.md |
Base system prompt |
config.txt |
Default settings |
This is the module map for structure, not a full inventory — src/agent8088/
also has browser_*.py (the browse_page engine), cli_anything.py,
fusion.py, mcp.py/mcp_server.py, model_catalog.py, ocr.py,
patching.py, routing.py, searxng_provision.py, task_runtime.py,
trajectory.py and a few more single-purpose modules.
Per user turn, run_agent() loops up to max_turns:
- Recall memory (outermost turn only) — the last genuine user turn is searched against the memory store and the best few facts are appended to this turn's system prompt. See Memory.
- Call the model with history + tool definitions.
- Strip reasoning from the content so chain-of-thought never leaks into the answer.
-
Find tool calls — native
tool_callsor the fine-tuned text-marker format, restricted to the allowed set. - No tool calls? Check whether the model tried to call something that doesn't exist. If so, tell it what went wrong and loop (bounded) so it can recover. Otherwise this is the final answer.
-
Deduplicate — an identical
(name, args)repeat is served from cached output instead of re-running, which breaks loops. -
Execute through
run_tool()→check_permission()→ the tool. -
Feed the result back, truncated to
max_tool_output_bytes, redacted, and wrapped if it came from outside. - Guard the answer before returning — system-prompt leak check, secret redaction, markup stripping.
- Capture memory (outermost turn only, after the answer) — one extraction call distils durable facts from what the user said and what was answered.
Steps 0 and 9 hang off run_agent() rather than the loop itself. The loop has
seven return points, so a capture hook inside it would be silently skipped by the
next one added; run_agent's finally cannot be escaped. It also already
distinguishes the outermost turn, which is the scope memory wants — a sub-agent is
handed a delegated task rather than something a human said, so it neither recalls
nor writes.
If the model backend errors mid-turn, the loop returns the best output it has rather than crashing the session.
execute_plan runs its steps in order and halts on the first failure, rather
than running the rest against a state the plan no longer describes. A step counts
as failed when the tool reported an error, or when its escalation went unanswered
or denied.
The result then says where it stopped and how many steps did not run, so a caller cannot mistake a half-executed plan for a finished one. Continuing past a failure is what turns one bad step into a cascade: every later step is built on an effect that never happened, and the transcript gives the failure no more weight than any other line.
A halted plan is not a rollback — steps that already ran stay run. The caller is expected to fix the cause and issue a new plan for the remaining work.
A step that reports success has only told you the tool did not raise. Setting
plan_audit=1 sends each mutating step to the readonly auditor sub-agent,
which inspects the real environment before the plan continues. A fail verdict
halts the plan on the same path as any other failure.
Mutating is decided per call, not per tool. A tool's mode names the
permission gate it goes through, and shell is the gate thirteen read-only
tools share — so testing the mode sent git_status, git_log, git_diff and
every ls to the auditor. Measured on glm-5.3, one git status bought a
thirteen-call auditor run, 111 seconds and 100% of that turn's tokens to
confirm that a read had read. _tool_call_mutates asks the question an audit
actually answers, and for a shell call it asks _shell_call_is_read_only,
which requires every &&/||/;/&-separated segment to clear the same
_readonly_shell the approval prompt uses. That relaxes sequencing only: a
command cannot be read-only enough to skip a confirmation prompt yet mutating
enough to need an auditor, and one write anywhere in the sequence condemns the
whole call.
This covers two paths, because a plan's work now happens on both. Steps inside
execute_plan are audited there. The tool calls an approved /plan makes are
audited in exec_tool, at depth 0 only — a sub-agent's writes are not the plan,
and auditing inside the auditor would set it verifying itself. That second path
was not always covered: plan mode used to force every mutation through
execute_plan, so hooking execute_plan was the same thing as hooking the plan.
Once approval started handing execution to ordinary tool calls, the default
/plan flow became the one flow with no verification at all.
An execute_plan step can declare acceptance; an ordinary tool call has nowhere
to put one, so the plan the user approved is passed as the criterion instead.
Without it the auditor grades "did this call take effect", which a write that did
confidently the wrong thing passes. With it, a write_file of hi against a plan
promising a revenue table comes back:
VERDICT: fail — the file contains only "hi", not a markdown table of quarterly revenue with at least four data rows — and the file is removed.
Verification is not free, so the cost stays visible. _exec_plan appends its share
to the plan's output; the post-approval path has no such output, so the share is
captured as the turn's budget is torn down and the CLI prints
verification cost this turn: N% of tokens. On a small turn that share is large —
a single audit call against one tiny write measured over 90% — which is the number
to look at before leaving auditing on.
Two deliberate asymmetries:
- A failed verification halts; an auditor that could not run does not. An inconclusive result is recorded in the output and the plan proceeds — a verifier that cannot reach the model must not be able to stop all work on its own.
- Auditing is skipped once the shared turn budget is spent. The auditor draws on
the same
_TurnBudgetas the work it checks (a fresh budget would be a free bypass), so it yields rather than starving the task it exists to protect.
Off by default: it costs one model call per mutating step. The case for turning it on is the unattended paths — gateway and cron — where nobody is watching and a half-done plan reported as finished is the expensive failure.
A step proposes a change; verification is what commits it. With
plan_audit_revert=1 (the default when auditing is on), a write_file step that
fails verification is restored to its exact pre-step bytes — created files are
removed, overwritten files are put back.
The boundaries are deliberate and worth knowing before you rely on it:
- Only the failed step is reverted. Steps that already passed verification stay committed; a failed audit is not a transaction rollback across the plan.
-
Only
write_filecan be undone.execute_shell,run_sandboxedandschedule_taskhave no inverse. Those steps reportnot reverted — no undorather than implying a rollback that did not happen. -
Large files are not snapshotted. Above
plan_audit_revert_max_bytesthe prior contents are not held in memory, so the write stands and the output says so. Declining to revert is acceptable; silently failing to revert is not.
The worst case is bounded by construction: a revert restores the exact bytes that
were there before the step, so a wrong fail verdict costs you the step, never
data that predates it.
_TurnBudget attributes tokens to the role that spent them (main, or
subagent:<type>), and each model-telemetry line carries the same role field.
After an audited plan the result reports the share, e.g.
Verification cost this turn: 22% of tokens (4400 of 20000).
This exists so the decision to run auditing is made from your own numbers. Published figures put auditors at roughly a fifth to a third of harness tokens; whether that holds for your workload is measurable rather than assumed.
Turning verification on multiplies a turn's model calls, so it brings its own
ceilings rather than relying on ones nobody set. max_turn_seconds,
max_turn_tokens and max_turn_cost_usd all default to 0, which means
disabled — on a default install _TurnBudget.exceeded() returns None after
two billion tokens, so "skip the audit if the turn budget is spent" never
fired. Three things bound it now:
-
plan_audit_max_per_turn(12) — verification calls one top-level turn may spend. Running out is not evidence a write was wrong, so it halts nothing and reverts nothing; it says so in the tool result. There is no unlimited value: to stop verifying, setplan_audit=0. -
plan_audit_timeout_seconds(120) — one auditor sub-run. -
subagent_timeout_seconds(300) — every other sub-agent._exec_subagentshares the parent's remaining tokens and cost, so delegating still cannot restart the turn's allowance, and a sub-run's clock is capped by whatever the turn itself has left — a 300s sub-agent cannot extend a turn with 60s on it. Before this atest-writerinheritedmax_seconds=0and ranfind / -maxdepth 5until the process was killed.
A sub-run's tokens are billed back to the parent by role when it finishes, so giving it its own budget bounds the cost without hiding it from the line above.
tools.txt is a pipe-delimited registry, not Python. A tool declares a mode,
and the mode determines its gating. Adding a row gives you a new tool with
correct permissions automatically — no new security code, and no way to
accidentally add a tool that bypasses the gate.
TOOL_SPECS (name → spec) and TOOLS_DEF (the JSON schema list sent to the
model) are built together from it. The test suite verifies their names match
exactly, so a registered tool can never be invisible to the model, or vice
versa.
Ordered outermost-first; each layer can only refuse, never grant:
1. allowed_paths — is this path in scope at all?
2. sensitive-file floor — credentials: refuse read AND write, always
3. shell-startup floor — refuse writes to .zshrc etc., always
4. write path zones — blocked / no-prompt / prompt
5. check_permission(mode) — readonly / full-auto / plan-only
6. shell classifier — safe inspection vs mutation, sees through sh -c
7. hard git blocks — shell push / reset --hard, always
8. SSRF guard — outbound URLs, including redirects
9. sandbox — OS isolation for whatever survived
10. output guards — secret redaction, untrusted wrapping, leak check
Layers 2, 3, 7 are the "always-on floor": no mode and no escalation grant
unlocks them. The git_push tool is the only way to push: it asks every time,
and an approval covers just that exact remote and branch. See Permissions & Security.
| Path | Contents |
|---|---|
~/.agent8088/config.txt |
Settings (0600) |
~/.agent8088/.env |
Secrets (0600) |
$AGENT8088_HOME/mcp.json |
User MCP servers (0600; defaults to ~/.agent8088/mcp.json) |
.agent8088/mcp.json |
Project MCP servers — override user-level |
~/.agent8088/gateway-sessions/ |
Per-chat history |
~/.agent8088/runtime/ |
Sandbox runtime |
USER.md |
Persona |
There is no database and no server process. State is plain files you can read and edit.
Mostly single-threaded and synchronous — deliberately, since it's an interactive tool. Three exceptions:
-
MCP client runs an asyncio loop on a dedicated background thread, with a
synchronous facade (
MCPRuntime) over it, so the rest of the engine stays sync. -
Gateway is async end to end (Slack Socket Mode,
discord.py), calling into the sync engine from the adapter layer. -
browse_pageruns oneasyncio.run()per call (the browsing library is async-only) alongside a short-lived filtering proxy on its own background thread. Both end with the tool call, so nothing async outlives it.
Honest list, for anyone extending this:
-
No cross-run cost accounting.
/usagereports session tokens andmax_turn_cost_usdbounds a single turn, but nothing aggregates spend across runs.model_telemetry=1writes per-call records to a local JSONL file; the gap is that nothing reads them back for aggregation. -
browse_page's reliability depends on the configured model. It can click, fill forms, navigate, and extract via natural-language instructions, but each step relies on the model returning strict, parseable JSON. Some models/providers occasionally wrap that JSON in markdown code fences or return malformed output; when that happens repeatedly within one task, the browsing agent gives up early rather than completing the task. This is a limitation of the underlying model's instruction-following, not ofbrowse_pageitself.For the same reason,
browse_pagenever asks the model for a screenshot: many providers Agent8088 supports are text-only, so vision is off unconditionally rather than assuming one particular model's capabilities. If a task stops early, try it again with a model that follows a strict output format reliably.
Source of truth: docs/wiki/ in the main repository. Edits here are overwritten by the next sync.
Start here
Guides
- Permissions and Security
- Sandboxing
- Model Providers
- MCP
- Messaging Gateway
- Skills and Subagents
- Memory
- Docker
Reference
Development