Repository navigation
Session Backed Agent Runtime Design #743
justintime4tea
started this conversation in
Ideas
Replies: 2 comments
|
This discussion is about Session Backed Agent Runtime Design. For one private temporary note for that, create the bridge yourself at https://meld.mergeinc.workers.dev. Say what it is for and what it is not for, and send the URL privately. It stays open 36 hours until the first reply. Each later reply resets 24 hours. The conversation stays on that same link. It is free, and there is no billing. It is not session identity, and it will not resume the old session. |
0 replies
Open questions:
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
The agent runtime — design overview
A guiding design doc to the 578 sub-issue proposals (posted later as discussions). Those are issues, each scoped to one change. This is the picture they add up to: what a session and a run are, the states they move through, who can touch them and how, and where the lines fall between the parts. Where a proposal/draft owns a detail, this points at it (posted later as discussion) rather than repeating it.
Revision 7 (2026-10-07): the sub-issue proposals are renumbered to follow their dependencies; the design is revision 6's (2026-10-05). Baseline:
origin/nightlyatcb8b5a4c, PR 719 at04074b90, PR 730 at69bfe5c4.One paragraph
The domain has four nouns, nested: an agent is a config entry; a session is that agent's identity over time — the thing a stream, history, and later memory hang off; a run is one unit of work in a session, from a prompt or headless start through every inference and tool turn to
[DONE]or another terminal state; a turn is one inference frame inside a run, rig's business and not addressable. Today a run is the HTTP request or A2A execution that started it: the request's task drives the agent, holds its handle, reads its events, and cancels it on disconnect. The runtime separates the two. A session becomes a thing the process owns, with at most one run live at a time; a run is started on its own task and identified by aRunId; a request becomes one of possibly several observers of the session. Everything the agent does, and everything that happens to its runs, is written to one ordered per-session journal asSessionEvents; every place that needs to show it — the SSE handler, the A2A executor, the CLI, a management endpoint — subscribes to that journal and projects it into whatever shape it serves. Commands go the other way, through the runtime: start, attach, detach, cancel, continue, resume. Nothing at a seam reaches into the agent.The shape
flowchart TB subgraph seams["Seams — project events out, send commands in"] direction LR chat["POST /v1/chat/completions<br/>→ OpenAI + aura.* SSE"] a2a["A2A executor<br/>→ task status + artifacts"] mgmt["/v1/sessions · /v1/runs<br/>→ SessionEvent JSON, summaries"] cli["standalone CLI<br/>→ terminal"] end subgraph runtime["Runtime — owns sessions and their runs (aura crate)"] direction LR registry["registry<br/>SessionId → SessionRecord<br/>RunId → SessionId"] journal["journal per session<br/>dense seq · replay · lag"] obs["observer set per session<br/>collecting / claiming · presence"] live["liveness policy<br/>Cancel / Continue / Park"] poll["approval pollers<br/>keyed by session"] end subgraph run["A run — one unit of work, many turns"] direction LR ctx["RunContext<br/>id · cancel · event channel<br/>tool-call queue · budget · nudge<br/>present · claimed"] agent["Agent<br/>one run of a PreparedAgent"] gate["HITL gate<br/>holds the run, reads presence<br/>awaits decisions on the bus"] end subgraph below["Below the run"] direction LR rig["rig<br/>one tool server per agent"] mcp["MCP clients"] park["park checkpoints"] store["session store<br/>runs · journals · approvals · leases · bus"] end seams -- "start · attach · detach<br/>cancel · continue · resume" --> registry journal -- "Stream<SessionEvent>" --> seams registry --> ctx ctx --> agent --> rig --> mcp agent -- "AgentEvent<br/>(via RunContext channel)" --> journal registry -- "LifecycleEvent" --> journal obs -. present · claimed .-> ctx live -. park .-> park registry -- "RunStore · JournalStore · LeaseStore" --> store gate -. approval decisions .-> storeTwo directions. Three things cross the seam boundary:
SessionEvents, read from a journal subscription. A seam never holds theAgentRun, never readsStreamItems, never subscribes to anything but the journal.POST /v1/approvals/{id}. The ingress writes it toApprovalStoreand publishes it on the bus; the gate, which subscribed when it routed the call, wakes and reads it (09). No channel runs from the web server into an agent. The runtime sees the outcome as events like any observer.The runtime is the one in-process facade. Seams call it and nothing below it; the HTTP API (08) is its public form. "Bus" in the sense of a transport is not the rule — the rule is that nothing under the runtime is reachable from a seam, and the only channel end a seam ever holds is the stream inside its subscription.
Vocabulary
AgentInfo::id): model, tools, MCP servers, prompts. The thing a session is of.chat_session_idtoday; A2A's context id). Owns the event stream, has many runs in sequence and at most one live, and is where history and memory attach.RunId, minted when it starts. A user's input enters a run only through HITL decisions and the turn nudge, between turns.PreparedAgent::prepare; serves runs one at a time throughbegin_run.Agent(the type)RunContextchildsharing id, observer, and cancellation.SessionEventAgentEvent(what the agent did) or aLifecycleEvent(what happened to the run or the session). PR 730 calls itRunEvent; the rebase renames it.SessionEvents, stored behind aJournalStorecapability of the session store (in memory today). The only reader of a run's channel; the source every subscription reads.Cancel(today's behavior),Continue, orPark. Chosen per start; a grace window may precede it.AgentStateSessionEvent → Option<frame>that a seam applies to its subscription. There is no projection trait; a projection is a function.A run's life
stateDiagram-v2 direction LR [*] --> Preparing: Runtime\:\:start Preparing --> Running: Started Preparing --> Failed: prepare / begin_run fails state Running { [*] --> Claimed Claimed --> Unclaimed: last claimant detaches<br/>ClaimsExhausted Unclaimed --> Claimed: a claimant attaches<br/>within the grace window } Running --> Finished: agent completes<br/>Finished { usage } Running --> Failed: agent errors<br/>Failed { error } Running --> Cancelled: cancel token fires<br/>Cancelled { reason } Running --> Parked: gated call parks under [hitl.park]<br/>or LivenessDecided(Park)<br/>Parked { checkpoint } Unclaimed --> Cancelled: LivenessDecided(Cancel) Unclaimed --> Unclaimed: LivenessDecided(Continue) Finished --> Retired: retention window elapses Failed --> Retired Cancelled --> Retired Parked --> Retired Retired --> [*] Parked --> [*]: /resume, or resume_on_start,<br/>starts a NEW run in the same session Finished --> [*]: POST /v1/runs with continues<br/>starts a NEW runNotes the diagram cannot carry:
Preparingis counted, andstartdoes not wait for it to end. The runtime holds a run from the momentRuntime::startis called, before the agent exists, so shutdown drains it ([BUG]: graceful shutdown ignores requests still in per-request MCP init #736's gap), andstartreturns the id then and there. On a warm session (10)Preparinglasts as long asbegin_run; on a cold one it includes MCP discovery and, with [FEATURE]: Atomic Fence for VFS Claims #581, the session claim. A seam that needsStartedattaches and waits for it; the record emits no events before it.Unclaimedis not a status. It isRunningwith an empty claim on the session and, if the policy has a grace window, a timer. The observable record of it is the pairClaimsExhausted→LivenessDecided { policy }in the stream.Runningto the run andBlocked { on }to the session (03's derivation). A run that stops because the agent asked a question (ClarificationNeeded) finishes; one that stops at an unattended approval parks. Both are answered by starting a new run in the session —POST /v1/runswithcontinuesfor the first,/resumefor the second — because a new user prompt cannot enter a run in progress ([FEATURE]: A run handle owning identity, events, and cancellation #625, [EPIC] Agent-Driver #211). HITL decisions and the turn nudge can, between turns, and do.AgentRundirectly and never goes through the runtime still gets today's behavior: drop the stream, the run cancels. The runtime's own cancel-on-unclaimed is the same behavior expressed as a policy.start.Parkon an agent with no checkpoint support isStartError::PolicyUnsupported, not a quiet downgrade toContinue.A session's state, as the endpoints report it
The session is what a product sees as "the agent", so its state is what
GET /v1/sessions/{id}answers. None of these is an event; each is constructed from the records 03 holds, in this order:AgentStatePreparingPreparingBlocked { Approval }pending_approvalsis non-emptyBlocked { ProviderRetry }RetryingRunningParkedParkedDoneIdle — booted but never asked anything — is
Donewith no last run. Failed and Cancelled are run outcomes onRunFacts, not session states; a session whose last run failed isDoneand will prepare again on the next start if its agent slot was lost.What each transition writes to the journal
Runtime::startsucceedsStarted { agent, prompt, timeout, liveness }ObserverAttached { observer }/ObserverDetached { observer, because: Released | Expired | Displaced }ClaimsExhaustedLivenessDecided { policy }Finished { usage }Cancelled { reason: Deadline | External | ClientTool | Unclaimed | Shutdown, message }Failed { error }Parked { checkpoint }ApprovalRouted { decision_id, route: Conversational | Webhook | Poll | Parked, because }Everything the agent itself does — text, reasoning, tool calls, progress, approvals, orchestration phases, provider retries once that signal exists, the final
Completed— arrives asSessionEventPayload::Agent(AgentEvent)through the run's channel, interleaved with the lifecycle in one sequence. Lifecycle events that belong to no run — an observer attaching to an idle session, the session claim — carry norun_id.Observers
A session has zero or more observers. Each is one of four things:
Rules that hold across every draft:
Released; detaching emitsObserverDetached; if that was the claim and a run is live,ClaimsExhaustedfollows. Only an attachment that crosses a process boundary — the relay in 08, the session claim in [FEATURE]: Atomic Fence for VFS Claims #581 — is a lease with a TTL, because only there can the holder die unseen; that is whereExpiredcomes from.How you interact
POST /v1/chat/completions)chatdefault (Cancel)ClaimsExhausted→ Cancel)message/send,message/stream)tasks/{id}:canceltasks/{id}:subscribe)Started(replays the journal)/v1/sessions,/v1/runs)POST /v1/runs— headless; liveness = theheadlessdefault (Continue); returns inPreparingGET /v1/sessions/{id}/events— collecting by default,?claim=true&presence=trueto take the helm;GET /v1/runs/{id}/eventsfor one runPOST /v1/runswithcontinues(a new run),POST /v1/runs/{id}/resume(a parked run)POST /v1/runs/{id}/cancelPOST /v1/approvals/{decision_id})A chat request end to end, once 05 lands:
sequenceDiagram participant C as Client participant H as chat handler participant R as Runtime participant J as Journal participant A as Agent (rig) C->>H: POST /v1/chat/completions (stream: true) H->>R: attach(session, Claiming, presence: true, from: Live) R->>J: Lifecycle(ObserverAttached) H->>R: start(StartRun { messages, headers, client tools, session, liveness: chat default }) R-->>H: RunId (Preparing) R->>A: prepare if cold → begin_run → Agent::stream(RunOptions) R->>J: Lifecycle(Started) loop until terminal A-->>J: AgentEvent via RunContext channel (the run's task appends it) J-->>H: SessionEvent H-->>C: project_sse → SSE frame → response body end alt agent completes A-->>J: Agent(Completed) R->>J: Lifecycle(Finished) J-->>H: SessionEvent(Finished) H-->>C: [DONE] H->>R: Subscription dropped → detach (Released) else client disconnects C-xH: connection drops H->>R: Subscription dropped → detach (Released) R->>J: Lifecycle(ObserverDetached), Lifecycle(ClaimsExhausted) R->>J: Lifecycle(LivenessDecided { Cancel }) R->>A: cancel token R->>J: Lifecycle(Cancelled { Unclaimed }) endAttaching to an agent that is already working — the case the epic exists for:
sequenceDiagram participant O as Operator participant M as GET /v1/sessions/{id}/events participant R as Runtime participant J as Journal Note over R,J: a headless run started earlier with liveness = Continue,<br/>its claimant long gone, still running O->>M: GET ?after=0&claim=true&presence=true M->>R: attach(session, Claiming, presence: true, from: Start) R->>J: Lifecycle(ObserverAttached { claiming, present }) J-->>M: replay: every SessionEvent still held, across runs, then live M-->>O: SessionEvent JSON, one per SSE data line Note over R: the session now has a present claimant - a gated call would take the conversational route (09)The bounds
Who owns what, and what may cross each line.
aura-events—agent.rs,run.rsAgentEvent,SessionEvent, the lifecycle enum,RunId,SessionId,Observer,DetachCause,LivenessPolicy,Liveness,RunCancelReason. Serde only; no agent, MCP, or provider dependency.aura::run_context,aura::builderRunContextandBoundRun(nightly, extended by PR 719),RunLease,PreparedAgent,Agent,AgentRun,RunOptions, hooks. The task-local scope. EmittingAgentEvents onto the run's channel.AgentRun.aura::runtime(new)RunId, draining on shutdown. Storage is the session store's; it holds the capabilities, not the data.aura::hitlRunContextit holds; awaits decisions through the approval store and the bus.aura::orchestration::parkParkpolicy,/resume, and resume-on-start call into it.aura-web-serverhandlers, A2A executor,aura-cliAgentRun. ReadStreamItems. Reach the MCP manager or a cancel token. Hold any channel end but the stream inside a subscription.aura::session_store,aura-web-server::session_storeRunStore/JournalStorecapabilities — in-memory impls now, durable later (#210 / #325) — and, with #581, theLeaseStorethat holds the instance's claim on a session and the relay's attachment to it. Each store bounds what it keeps.Four consequences of drawing the lines there:
auracrate, not the web server, so a process with no HTTP layer still has sessions, journals, and observers — and stops building server state in-process to round-trip through SSE.SessionEventand returns a frame. Parity tests (05, 06) drive a scripted run through the old handler and the runtime-backed one and compare frames byte for byte.RunContextchannel, the orchestrator'sStreamItemchannel,ToolCallObserver), its receiver is threaded through three struct layers, and the web server holds a oneshot into the HITL gate. After 02, 03, 05, 09, and [FEATURE]: Remove the request-scoped broker path #627, none of that remains.Invariants
append. A gap is detectable; a relay or projection never renumbers. A run boundary is a position in the stream, not a second stream.Gap; the run and its other observers never wait.begin_run, because rig's tool server is agent-scoped and two concurrent runs would misattribute tool events silently.AURA_CUSTOM_EVENTSsays; a client that cannot see one cannot answer it.Persistence and resumability
"Resume" means four different things here, and only one of them exists today.
ParkedRuninorchestration::park, versioned (parked_run_v1.json), files0600(#683), keyed byrun_owner_id = run:{run_id}/resume(08) and resume-on-start (07) are the entrypoints it lacksJournalStoreread by session id plusRunStoreby session; the agent slot prepares again; a parked last run resumes by policyThe last row is the constraint that shapes the others: the unit of resumability is the run boundary, never the middle of a turn. A run ends — finished, parked, cancelled, failed — and what comes next is a new run in the same session, seeded from something durable: the session's stream or stored history, a park checkpoint, or the previous run's answer. Nothing in the runtime tries to pick up a
Runningrun. That is whycontinues,/resume, and resume-on-start all start new runs, and why a session's stream never has two writers at once.What is persisted, by structure
SessionEventand everything inside itJournalStoreon the session store, keyed by session — the in-memory impl now, a durable impl later (#210 / #325)ParkedRun,CheckpointRefCheckpointRefis the pointerRunFacts— a run's status, terminal event, usage snapshot, timestamps, session, agent, first and latest sequence, checkpoint ref, pending decision ids, lineageRunStoreon the session store, besideApprovalStore,TaskStore,SkillInvocationStore— in-memory impl now, durable laterRunRecord; what 08'sRunSummaryis built from; each field is a fold over the session streamParkedApproval, decision records)ApprovalStore, with the bus topicapproval:{decision_id}announcing a writeTaskStore(upstream trait)run_idandsession_idso a resubscribe elsewhere can find the stream and the cursorLeaseStoreon the session storeRunContext,AgentRun,PreparedAgent, hooksObserver,ObserverSet,SubscriptionObserverrides in the journal as historySessionJournal,Cursor,JournalItemJournalStore; the struct, neverSessionEvents leave itAgentState,BlockReasonPreparedAgentCache,CacheKey,HeadersFingerprintLiveness,LivenessPolicyRunFacts, inStarted, and on the wiregracemust serialize in milliseconds like every other durationStartRunRequest,SessionSummary,RunSummary, …)RunSummaryisRunFactsplus instance-local observer counts;SessionSummaryis constructedThe gap: session history
POST /v1/runsinto an existing session — a headless start, or the next turn withcontinues— implies the server can reconstruct the conversation so far. Today it cannot:/v1/chat/completionsis stateless and the client sends the full history every turn; only the CLI persists conversations, locally. A parked run is fine — the checkpoint carries each worker's conversation — but a run that finished with a clarifying question has had no server-side memory of the exchange.Keying the stream by session changes what is possible without deciding what is done.
Started { prompt }and each run'sCompletedcontent are the user and assistant turns in order, so history is derivable from one stream load; skill invocations are already persisted per session and spliced back (e54d2bc6). Whether the server reconstructs history from the stream, or #210 stores messages beside it, is #210's decision, and it is what "threaded history under the session" asks for. Until it is made, the drafts keep today's contract:StartRunRequestcarries the messages, and the runtime seeds the new run from them. What must not happen is the runtime quietly rebuilding history from a bounded in-memory journal and calling it the session's.The journal, shaped for durability now
A durable journal is later work (#210 / #325), but the shape that makes it a second implementation rather than a refactor costs nothing today, so the drafts commit to it. Five decisions, each of which would be a rewrite if taken the other way first:
Storage is a trait on the session store, keyed by session.
JournalStore(events) andRunStore(facts) joinApprovalStore,TaskStore,SkillInvocationStore, andEventBusas capabilities ofSessionStore, with in-memory implementations that are today's behavior. The runtime and the journal holdArc<dyn …>. This is the pattern the codebase uses for every other piece of cross-request state, ande54d2bc6([FEATURE]: Skills invocation persistence and rehydration #396) is a complete worked example: one trait, memory/file/Redis backends, a record with aversionfield probed before the full parse so a foreign-version record is skipped rather than mis-parsed, a per-session cap in memory, TTL where the backend has it, and file names derived from a v5 UUID of the session id so a client-supplied id cannot address outside the store. A durable journal is that shape with the same key. The alternative — a concrete journal whoseVecDequeis the storage — means durability rewrites the struct and every place that assumed the ring buffer's semantics.One writer per session at a time, through one serialized
append. Every event of a session — the agent's, via the channel PR 738 provides and the live run's task, and the runtime's lifecycle, appended directly — goes through oneappendthat takes the journal's lock, mintsseq, awaits the store, broadcasts, and releases.Startedis the run task's first append, before the agent produces anything; the terminal event is its last, after the agent's channel closes. In one process the lock and the one-live-run rule make the writer unique; across instances the session claim does. The store therefore sees each session as a strictly increasing sequence from one writer at a time. Two unserialized appenders would give a store out-of-order arrival the momentappendawaits; a side channel from the runtime into the run task would serialize them at the cost of one more channel end and a task that has to exist to carry it, which is the shape the design review asked us not to grow.Subscribe joins the broadcast first, then reads the store. A subscriber attaches to the live broadcast, notes the first live
seqit sees, reads the store up to one before that, and drops any live event whoseseqis at or below the last one it replayed. Sequence numbers make the dedupe exact, and no lock is held across the read. The in-memory version could instead snapshot the ring under a lock and subscribe in the same critical section — simpler, and impossible once the read awaits a database. Writing the async-safe algorithm now means the durable store slots in with no change tosubscribe.Gaphas one meaning. A subscriber missed events, andresumed_atis the first it will see. Two things cause it — a lagged broadcast receiver, or a cursor older than the store still holds (ReadError::Truncated) — and the consumer cannot tell them apart, nor needs to. What must never exist is a second kind of hole: a store that elides text deltas to save space would produce gaps that are not lag, and every projection would have to learn the difference. Every event is stored. A run's delta volume is bounded by the turn count; a session's by the store's cap.The terminal write is one store call. When a run ends, its last journal event and its final
RunFactsgo to the store together throughRunStore::finish(terminal, facts), so a backend that supports transactions can make them atomic. Two separate calls in the in-memory version would force every durable backend to accept a journal whose last event isFinishedand no record — or a record and no terminal event — after a crash between them.The one thing durability changes that shape cannot hide is back-pressure: a slow store slows
append, a slowappendfills the run's channel, and a full channel makes the agent wait onemit. That is correct — a run should not outpace what its journal can record — and it is a store-tuning question, not a design one. In memory it never happens.What durability does not need from these structures: a schema version on
SessionEvent(the store versions its documents, as the park store does), a way to merge two writers (there is never more than one), or a way to rebuildRunFactsfrom the journal (the facts are written, not derived — though they could be).Rules for the serializable types
These are the footguns. Each is either already avoided by PR 730 or is a one-line decision to take on its rebase.
RunIdis a UUID. HITL already parses the orchestration run id as one (orchestrator.rs: "orchestration run id … is not a valid UUID") and the park owner key isrun:{run_id}. ARunIdminted as anything else breaks both the moment 03 unifies the ids. UUIDv7 sorts by time, whichGET /v1/runswants.SessionIdis client-supplied and stays a string; every store keys it through a v5 UUID, as the skill store already does, so a client cannot address outside the store.LifecycleEventuses#[serde(flatten)](Finished,Cancelled) and every enum in the family is internally tagged. Neither works withbincodeorpostcard. JSON and MessagePack are fine — andsession-storage.md§13 already leans JSON. Write it down so nobody reaches for bincode "for speed".kind; its variants' payloads are internally tagged withtype(andRunCancelReasonflattens areasontag intoCancelled). PR 730 hit this once —Observer.kindcollided with the envelope'skindand had to be nested — and the roundtrip test caught it. Every future lifecycle variant with a field namedkind,type, orreasonis the same bug. Two ways out: a test that asserts no variant field uses a tag name, or adjacent tagging on the envelope (#[serde(tag = "kind", content = "event")]), which puts the payload one level down and removes the class of collision. It costs one nesting level on the wire. The envelope has no producer yet, so this is the moment.#[non_exhaustive]does nothing at deserialize time. It stops downstream exhaustive matches; it does not make an old binary tolerate a new variant. A persisted journal replayed after a rollback fails on the first unknown event. Tolerate it at the store boundary — an entry that will not deserialize becomesJournalItem::Unknown { seq, raw }— rather than in the type.SessionEventis a wire type; the park precedent (parked_run_v1.json) versions the stored document. A durable journal wraps its entries in a versioned record and migrates there.#[serde(default)]on added fields covers the additive case, asStarted.livenessdoes.duration_msmodule is the convention (Deadline { after_ms },Started { timeout_ms }). ADurationwith plain serde derives becomes{ secs, nanos }—Liveness.gracemust not.seq, never byat.Timestampis wall-clock milliseconds and exists for display. A session's sequence is written by one writer at a time — the lock in a process, the claim across them — soseqis total and dense andatneed not be, even across a claim moving between instances.SessionRecordandRunRecordhold aJoinHandle, locks, atomics, anArc<RunContext>, and a prepared agent. Neither may ever grow aSerialize. The facts a store wants live in a separateRunFactsthat does, and the run record holds one; the session's outward state is constructed intoAgentStateon each read. Then "what can be persisted" is a type, not a convention.Started { prompt }and everyTextDeltaare user content. A persisted journal gets the treatment Webhook poll delivery (1/4): config surface, pending-approval scan, owner-only store files #683 gave park documents: owner-only files, or a store with equivalent access control.What is real and what is planned
AgentRun,RunOptions, hooks, task-local run scopec530ee4e)RunContexte5522888)SessionStore::skills()— the capability pattern, versioned recorde54d2bc6, #396)PreparedAgent/Agent,RunLease,RunContextwith tool state andchildRunEventkeyed by runSessionEvent;RunContext::idas theRunIdnewtype;DetachCause;RunCancelReasonopenedStreamItemchannel andToolCallObserverretiredDeliberately later
Cross-instance subscription and claim (#581, #210/#325) — the
LeaseStoreand the relay. A durable journal. Server-side session history (#210 decides whether it is derived from the stream or stored). Agent memory under the session, whose write claim rides on the session claim. OTEL as an observer (spans are set inline today with context an event stream does not carry). Observer plugins, cron and webhook triggers — 04's model and 08'sPOST /v1/runsare the seams they would use. Mid-run steering by a new user prompt — waits on #211 dropping rig.All reactions