RFC: Foreground/background sessions with a durable inbox #5231
Replies: 2 comments 1 reply
|
One design fork worth making explicit before the daemon ships. The inbox is discovered (CDC from SCM and worker facts), so every item then needs the system to infer what the operator meant by it. I ran that shape, a background coordinator that found its own work, classified it and asked when unsure, and the questions were the failure. They went to a file nobody opened, thirteen piled up, and every task behind one stopped. The daemon also rode an ambient credential that expired silently, so it failed dark for five weeks while its alarms fired into channels only the dead system read. The inversion that fixed it is that work reaches the agent when the operator marks it delegated, and marking runs a short intake while the intent is still in the operator's head (what done looks like, which paths may be written, how the two foreseeable forks should be ruled). A discovered event can still raise a card. It just cannot be a mandate. Your dispatch authority and approval gates put the human at the end of the pipeline, and the intake puts them at the start, where the ambiguity is cheapest to remove. The durable continuation obligation is the piece I'd keep exactly as written. I maintain a reference architecture that documents the queue pattern at https://github.com/jimy-r/agent-workspace-architecture/blob/main/PATTERNS.md#14-delegation-is-a-queue-you-fill-not-work-the-agent-finds |
|
This discussion is about RFC: Foreground/background sessions with a durable inbox. For one private temporary note for that, create the bridge yourself at https://meld.mergeinc.workers.dev. Say what it is for and what it is not for, and send the URL privately. It stays open 36 hours until the first reply. Each later reply resets 24 hours. The conversation stays on that same link. It is free, and there is no billing. It is not session identity, and it will not resume the old session. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I propose two orchestrator sessions: foreground for chat, planning and delegation; background for tracking workers and handing off already-authorized tasks.
Example: backend API verified -> frontend worker unblocked -> test worker checks the combined app. New scope decisions return to FG.
Goal: less chat contention, less repeated context, faster handoffs while you're away. Separate contexts; the daemon checks grants/revisions and handles recovery.
Problems this proposal aims to fix
These are planned improvements, not features already shipped.
Proposal in brief
Split orchestration into two model sessions:
Targets: user requests competing with coordination work, repeated context injection and delayed handoffs. Existing worker CI/review automation stays with workers; this proposal changes orchestration, routing and delivery, not the existence of auto-fix.
Example: "Build email signup: backend API, frontend form, then test them together. Keep it local; do not push." FG records the plan and delegates. BG can hand a verified API contract to the frontend worker and request a test worker for the exact combined artifact once prerequisites pass. Merely finishing both branches is insufficient: the integration worker must assemble the intended revisions in an isolated workspace and attach test evidence to that combined revision. Missing evidence keeps the dependency pending.
Context: FG retains user conversation, requirements, decisions and relevant progress. BG receives selected instructions, the current task contract, worker state, dependencies and evidence references. Neither receives every tool log by default. The contexts are independent; durable facts and scoped grants are authoritative, not agreement between model summaries.
Communication: daemon code reads/writes SQLite on the models' behalf. Conversation/intent records, evidence, inbox updates, consumer progress and delivery manifests have separate roles. The durable inbox is a queue with priorities, batching and retention, not necessarily one messages table or strict FIFO. Background processing does not consume foreground delivery. A controller reads relevant pending revisions and composes model input; a DB poll alone does not start inference. Reads do not delete records; retention follows unresolved references and receipts. Compaction starts a new context epoch with a bounded reseed. Existing SQLite, Chat queues and trigger CDC are reused.
Interactive examples and component state
Smallest implementation
Reuse the Go daemon, SQLite, Chat controllers and worker executor. Add one BG controller, a bounded inbox selector and one managed-action drain. No separate queue service or general workflow engine.
The work queue contains authorized operations; the inbox contains observations. Events may change eligibility or raise a suggestion, but cannot put unrelated work into the queue. Deterministic checks handle known transitions; BG inference handles judgment within the delegated task. With no new relevant input or due investigation, no model call is needed.
Incorporating jimy-r's feedback
His comment reports hidden clarification questions and an expired credential that left his earlier coordinator failing silently. His linked project is a reference architecture, not an AO fork. Its successor uses explicit delegation, short intake, questions on the task card and an on-demand drain; it deliberately gives up unattended progress. Pattern and tradeoff.
done_when, context, limits and foreseeable decisions upfront. Keep questions visible on their task; block only affected work when scope is known.done_when, limits execution to an on-demand session, stages outward actions, and uses timestamp claims with a 24-hour convention. AO uses existing user authorization, revision checks and durable atomic claims; an expired claim is not proof a worker stopped. Queue skill sample.The public repo documents patterns and skill/script samples. Its reported first successful drain is useful experience, not validation of AO's crash, concurrency or unattended guarantees.
Runtime and ownership
One controller/writer per provider conversation. The background has its own persistent conversation and bounded execution budget. Foreground and background inference can overlap; neither waits for the other's model call or investigation. Keep shared SQLite transactions short; inference and provider I/O run outside them.
Key background ownership to the durable foreground project conversation, with a distinct subordinate runtime identity and one active ownership generation. An ordinary
KindOrchestratorstartup rebinds the project narrative; it must not be used to create the background conversation. Foreground replacement transfers/rebinds the background owner and fences the old generation before new dispatch. Existing project ownership.Lifecycle retains deterministic eligibility: CI/review opt-ins, stack dependencies, conflict rearming and session completion rules. At the
pendingNudgeboundary, submit a typedReactionIntentwith cause, evidence revision, target, policy and operation. The dispatcher owns managed sends; remove the overlapping direct-send path for those causes. Intercepting only the final textSendloses this metadata. Include PR reactions and internal review-batch delivery in this routing audit. Current reaction boundary.Audit
ApplyTrackerFactstoo. It has no production caller in this snapshot; any wiring must submit the same managed intents. This does not add a new tracker integration.A routine CI repair can go straight to the authorized worker through deterministic eligibility; BG inference is not mandatory for every CI failure. BG can advance pre-authorized dependencies and handoffs. Foreground involvement is for new plan changes, cross-worker decisions outside that plan and user interaction. Worker-action urgency and user-interruption urgency are separate fields.
Internal review delivery needs an asynchronous receipt contract. Persist each review run's operation link when enqueueing;
enqueuedmust not meanReviewDeliverySent. Update delivery state from supported executor evidence, retaining unknown outcomes. Recheck target SHA and opt-in at dispatch. Alias each internal run'sGithubReviewIDto its SCM observation and track batch membership so the same review cannot start two fixes. Current synchronous caller.Dependent work keeps the durable continuation guarantee inside the pending
managed_action: prerequisite IDs, required outcomes/artifact revisions and authorizing task/grant revisions. No separate continuation table or lifecycle. Verify the completion predicate before atomically claiming that same operation. Idle, session termination and Chat turn completion alone do not prove task success. Unknown prerequisite outcomes leave the dependent action pending. An already-submitted action retains its delivery state, including acceptance uncertainty. Turn-state semantics.Reuse the session manager/existing worker executor. No second unrestricted agent editing the same worktree. Human commands remain available and supersede affected pending intents through the same resource controls.
Task intake and clarification
An explicit user request to do/delegate work admits the task; a discovered event does not. Intake drafts a concrete next action, existing project/context links, checkable
done_when, budget and foreseeable fallback decisions from the instructions. Ask only for missing essentials; no mandatory questionnaire or repeated approval. Persist a versioned task contract linked to those messages and one task-scoped grant. The contract describes the goal, completion criteria and constraints; the grant alone holds allowed operations, write roots and authorization limits. Dependency predicates live on pending actions. The task references the grant ID instead of copying its permission fields or pinning its revision; actions carry the revision they validated.Trusted daemon intake records only authority supported by the human instruction/current policy; model prose cannot mint or widen a grant. Narrow the grant to the delegated task, even if the project policy allows more. At dispatch, check current task/intent, grant and policy revisions, plan constraints, operation, write roots, budgets and prerequisites. Permission changes revise the grant; plan changes revise the task. Either fences affected unaccepted actions until re-evaluation under current authority; preserve the operation ID and any existing delivery uncertainty. Existing deterministic lifecycle automation binds its typed rule/opt-in to the existing rule-scoped policy grant, without fabricating a task. Model-proposed actions must reference a task-scoped grant and cannot claim that exception.
Discovered events can update an already-delegated task or propose a card/notification. A broad project permission alone is not a mandate for a new task. A new plan decision goes to FG, and a permission expansion requires the human. Known decisions covered by the delegated plan continue without asking again.
Clarifications are durable task records, not a file the operator must discover: question ID, affected task/dependencies, recipient, evidence/intent revisions, blocking reason, first-seen time, answer and disposition. Surface them on the task and in the orchestrator's pending-decision view. Dedupe repeated questions, preserve age, and resurface stale unanswered blockers on return. An answer updates the contract/revision and revalidates pending dispatch; it is not itself a provider approval receipt. Never treat silence as consent.
Hold only the affected task and dependents when scope is known. Other authorized tasks may continue. Unknown project-wide intent can require a wider hold. A provider-level input/approval block on the BG runtime still stops that runtime's inference; deterministic ingestion and already-independent authorized work continue. Do not create a second BG writer or auto-answer the runtime to bypass the block.
State and ingestion
Logical records, reusing existing tables where their contracts fit:
Persist
context_epochand reseed status on each existing Chat conversation record; checkpoints/manifests reference that epoch. Restart must not reset it or treat a prepared reseed as provider-accepted.triage_statetracks BG processing; FG delivery is derived from accepted manifest entries for the exact update revision in the current conversation/epoch. Retain updates while either consumer still needs them. Question dedupe uses task + decision identity, preserving first-seen age; changed question/evidence revisions require revalidation.Grant scope is task or deterministic rule, never an unscoped fallback. Retain operation-keyed attempt history with immutable envelope references, using existing outbound storage where it fits; persist state transitions and receipts. Each transport attempt has its own ID and exact envelope reference; retries retain the operation ID.
context_manifestseparately records FG turn/steer context delivery.Capture before overwrite. Persist significant observed transitions and immutable evidence with the source write, before check reruns or comment edits replace mutable rows. DB triggers remain the CDC producer. Asynchronous hydration of current rows cannot recover old logs. Retain evidence until consumer processing and while referenced by unresolved items; use bounded excerpts/blob limits and make truncation visible.
Source IDs, fingerprints and ordering are distinct. Prefer provider run/edit versions when available; otherwise record a local monotonic observation revision and reconcile current upstream facts. A hash is neither a run ID nor proof of source order. Preserve captured
fail -> pass -> same failedges; collapse repeated unchanged snapshots. Partial fetches remain unknown, never “all resolved.”Add a new migration covering
review_hashandreview_partialchanges in the PR CDC trigger. Cover comment insert/update/delete paths with compatible filtered invalidation triggers and evidence capture; preserve existing payload contracts. Include semantic worker state, blockers and outcomes. Filter streaming churn before consuming the triage budget. Current PR trigger, mutable check upsert.Persist inbox changes and consumer progress atomically. The live CDC callback is only a wake hint; it never runs inference. Resume the durable cursor, reconcile on startup and periodically, and record history loss explicitly. Update the deletion watermark transactionally with all relevant pruning/deletion paths; read it consistently with replay. Epoch changes also force reconciliation.
MIN(seq) > cursoris not a valid gap test. Live poller, retention queries.Reconcile on
epoch mismatch OR max_deleted_seq > cursor.seq; the watermark conservatively includes interior deletions. Use a consistent local snapshot withH = max(retained_head, max_deleted_seq). Preserve surviving captured observations and reconcile current facts into the inbox before committing(epoch, H). An empty log still uses the watermark; never reset to zero merely because no rows remain. Deletions after that snapshot are checked on the next pass. Remote refreshes and inference run outside this transaction.Exclude the background runner's own session/conversation events from model triage; observe its liveness and approvals through deterministic health handling. Foreground activity-only churn is not a new task event. Keep explicit human-intent sync, approvals, verified outcomes and worker semantic changes through their designated routes. This prevents
background turn -> activity CDC -> background turnfeedback without discarding worker results.Background continuity and tools
Sync selected human/assistant messages, durable scoped intent, verified tool outcomes and pending action/approval state into the background. Exclude raw tool chatter by default; retain evidence references for retrieval. Assistant prose and generated summaries are claims, not new evidence. Tag deliveries with origin/cause IDs so an injected digest does not return as a fresh event.
Each decision names its input revisions. Reject stale classification commits and superseded fetch results. At command acceptance, check intent/control and source revisions again. Increment the human-input revision atomically with enqueue, edit, withdrawal and order changes; fence affected automation until intent synchronization catches up with the surviving instructions. While scope is unknown, hold that orchestrator's pending automation. Explicit pause/cancel changes scoped control state immediately. Existing queue mutations.
Checkpoint the sync cursor, unresolved work and evidence references. On restart/provider-session loss, rebuild a bounded context from durable state. Compaction preserves outstanding obligations and their provenance; an indefinitely growing transcript is unnecessary.
The background gets enforced read/evidence tools plus typed daemon commands for authorized repair, review, continuation and worker creation. Validate operation, scope, target, concurrency and budget outside the model. Execution uses the worker's existing permission model. No arbitrary shell, filesystem writes or raw
ao sendbypass in the background runner.Add an enforced
ToolPolicystartup contract and corresponding capability: scoped read roots, exact allowed tools, no inherited provider/MCP tools, and unattended execution. An allowlisted MCP server alone is insufficient. Admission requires an adapter implementation and isolation tests; a capability flag alone is not enforcement. Refuse unsafe profiles withisolation_unavailablehealth status. Deny forbidden tools at the runner/tool host before execution; post-execution events cannot enforce this. Unexpected background approval/input blocks halt triage and raisetriage_blocked; never auto-answer them. Deterministic ingestion/deadlines/health continue. Current additive configuration.Bound batch size and wait, investigation duration, calls, tokens, retries and wake frequency. Delegate long investigations through existing workers. At capacity, coalesce/retain pending metadata and expose backpressure; never silently mark work handled. Reserve foreground capacity in local scheduling; shared account quota/outages can still constrain both sessions. Deadlines and ingestion/backlog health checks run without the classifier.
Context limits and usage
The main release needs per-adapter FG/BG context limits, not only limits on injected inbox deltas. Account for system/tool definitions, retained conversation, selected evidence, output reserve and a safety margin. Before submission, if the estimated request plus reserves would exceed the configured session limit, compact at a supported safe boundary or rebuild a bounded context through a supported rollover path. Preserve current intent/grants, pending questions, dependencies, operation IDs and evidence references. Start a new context epoch and reseed after compaction; never replay actions to reconstruct narrative.
Compaction/rollover is serialized by the owning controller. Queue edits and human input survive. Before a planned runtime replacement, reconcile accepted turns and unknown sends. Unexpected runtime loss instead carries unknown sends forward under the same operation IDs; recovery can proceed without resending them. Transfer logical ownership and fence the old generation before resuming managed dispatch. Expose a Compact control through supported adapters. If safe compaction/rollover or reliable context accounting is unavailable, expose that limitation and hold affected automated submissions rather than claim a hard bound or silently cancel/restart an active turn. Keep human messages durable and the UI available. If a human turn cannot fit or its fit cannot be established, show
context_recovery_requiredwith the retained message and supported recovery options; do not silently submit it into an unbounded context. Resume that same queued input after recovery and revalidation; report provider context errors explicitly.Expose per-turn and aggregate input, cached input, cache-write input where provided, output and context estimates, split by FG/BG/workers. Missing provider measurements are unknown, not zero; do not infer subscription usage percentages directly from API token prices. Preserve stable prefixes when possible. Do not send keepalive model turns just to maintain a cache. Cache expiry can still make the next turn expensive; compaction itself also costs tokens. Numeric context limits, margins, batch waits and reconciliation/health intervals are adapter/deployment configuration that must be selected and tested before release; this RFC does not claim tuned defaults or measured savings.
Independent health monitoring
In addition to deterministic checks inside the daemon, ship an independent local watchdog/supervisor for unattended operation. It runs outside the daemon and model processes, with an explicit OS-supervised lifecycle; desktop window closure must not silently stop the watchdog while unattended work is enabled. Its monitoring state stays under AO's data directory. No new network listener is required.
At unattended admission and before each new unattended dispatch, require a fresh watchdog heartbeat. Missing/stale health holds new unattended dispatch and surfaces
watchdog_unavailable; already accepted work is not cancelled. The watchdog retains incidents in its own local store under AO's data directory, outside the daemon database's write path; the daemon imports/reconciles them by incident ID after recovery.The watchdog also writes its generation and heartbeat timestamp to that store. The daemon reads it without writing, outside its SQLite transaction, and checks freshness again at dispatch against the configured bound. Missing, stale or invalid heartbeats fail closed; a cached successful read must not extend freshness.
Reuse AO's existing supervision and readiness facts where they fit. The daemon publishes a small health snapshot: last successful ingestion, provider readiness/auth failures, and task progress against expected state/deadlines. The outside watchdog checks process liveness and snapshot freshness, and delivers incidents; it needs no model or second provider integration. Missing/stale snapshots are failures, not healthy defaults. A live process is not proof of useful progress; an idle project is not a stalled task. Keep observer success distinct from task progress. Use known credential expiry only when exposed, without extracting credentials. Bound probe retries and exclude health signals from model feedback loops. Existing desktop supervision alone does not provide this unattended monitor.
The watchdog must notify through a local OS/desktop path independent of the failed daemon, its model, and the provider credential. Retain a visible health incident for the next app open, with last success and affected work. Test daemon death and expired auth with the ordinary daemon notification path unavailable. Detect expected shutdown/suspend and recheck health on resume; don't report an intentional stop as a crash. No local monitor can deliver while the machine is powered off, and no unattended guarantee applies if both watchdog and daemon are unavailable. Persistent OS supervision/restart and missing-watchdog checks on app start mitigate that limit; external/off-device delivery is outside this release. Faults never justify blind task replay.
Foreground delivery
Known-condition, deadline and health alerts bypass model triage. Semantic urgency still depends on classification and can be missed. Unsupported steering must have a visible fallback; do not simulate it with cancel-and-resend. Cancellation is a separate operation. Preserve guarded TUI worker execution with honest delivery status; automatic TUI foreground context insertion is outside this release.
Add a composition hook inside Chat dispatch, after selecting the human request and before provider submission. Routine context and wake eligibility stay in the inbox. Under the controller's dispatch lock, choose an eligible human turn first; only an idle controller with no eligible human turn may claim one coalesced wake and create its automation turn. A human request arriving after that claim cannot retroactively preempt an accepted wake. Recheck scope, source, approval/input state, generations, failure hold and resource ownership. Freeze a bounded manifest and exact payload durably before provider I/O. Omit stale optional context without dropping the human request. Current drain, failure gate.
Managed urgent guidance uses the controller's direct
Steerpath, extended with preflight and durable attempts; queued-turn promotion remains human-only. Prepersist the payload/manifest against the existing active turn and targeted provider turn, then link its activity receipt when available. No second turn is created. Capability absence, no active/steerable turn, non-steerable turn and unsupported content trigger notification plus deferred delivery. Ambiguous transport or post-acceptance persistence failures remain uncertain, without blind re-steering. Current steer behavior.Use a prompt-composition boundary that preserves the human's display text separately from the actual provider payload. Include constrained classifications, selected evidence and source references within a fixed context budget. Label external text as untrusted; neither labels nor summaries grant authority. Test the actual serialization on each supported adapter: Codex currently sends
msg.Text, so silently relying onmsg.Contentwould lose context. Codex serialization.Reuse Chat turn/message storage; add exact revision manifests and payload references. Provider IDs and generation success are not universal acceptance receipts. Uncertain acceptance stays uncertain until reconciled.
After conversation rollback/fork, handoff, compaction or provider-session loss, increment that conversation's
context_epochand build a bounded reseed from current durable state. Prior manifests remain historical receipts; they no longer establish what the new context contains. Include current intent/task/grant revisions, blocking questions, pending action prerequisites and relevant evidence references. Record the reseed manifest and its actual acceptance in the new epoch. Within an epoch, manifests still track exactly which update revisions were submitted and accepted. Context epochs are separate from runtime ownership generations and CDC epochs.If required control state cannot fit, use the visible context-recovery hold; do not mark the reseed complete. Optional facts that do not fit remain pending. Planned replacement waits for send reconciliation; after unexpected loss, retain
acceptance_unknownand reconcile by observation before any resend. Missing narrative never authorizes action replay. This avoids reconstructing exact evidence survival across provider summaries.Command acceptance, cancellation and alerts
Persisted acceptance defines the ordering with pause/cancel. Reject stale unaccepted commands; request supported cancellation of accepted work and report what remains running. A gateway cannot undo completed effects or fence every later shell command of an unrestricted worker. Serialize conflicting managed work by worktree/session/PR as appropriate; use separate worktrees for concurrent edits.
Use stable idempotency keys through services that support them. A retry reuses its operation ID. Source/intent revisions are preconditions, not permission to repeat an unresolved operation. Reconcile existing ownership before accepting a successor for the same cause. Suppression, exhausted retry budget, accepted delivery and resolved work are distinct outcomes.
Persist a submission attempt before transport I/O. Delivery state distinguishes
not_attempted,queued,submitting,acceptedandacceptance_unknown; operation outcome and source resolution remain separate. A guard'sSentplus error, unacknowledged TUI write or abandonedsubmittingattempt is uncertain. Guard refusal permits later re-evaluation; uncertain attempts suppress automatic replay until reconciliation. Source recovery may remove the need for work, but does not prove the earlier message was received. Guard error semantics.Extend the existing executor interface to carry the typed command envelope through native Chat queues and interface-transition outboxes: operation ID, source/intent/grant revisions, resource claim and target generation. Every drain revalidates before provider I/O; cancelled or superseded envelopes cannot become plain-text sends. The current
Sendcan return after enqueueing, sonilis not a provider receipt. Persist queue ownership separately from actual submission/acceptance evidence. Current queued routing.Ship the dedupe fix for the verified key collision:
comment:<PR URL>:<comment ID>instead of one shared PR key. Include body changes in observed revisions and re-evaluate material edits without blindly reinjecting every edit. Old shared keys cannot reconstruct per-comment receipts; document possible one-time repeats during migration. Key collision.Add cause-specific notification identity and resolution, including schema, unresolved-list/count queries, API types and UI. Two urgent causes on one PR must remain independent. Seen does not mean resolved. Use a daemon/project health surface for ingestion or provider failure without inventing a session. Show pending age, evidence, action disposition and the reason for a wake. Repeated polls must not reset deadlines. Existing notification assumptions.
Implementation plan: one release
These are implementation slices, not separately shipped passive/active versions. Shadow runs and feature flags are validation/cutover tools. For each managed cause, switch ownership atomically between old and new routing; never enable both dispatch paths.
Release gates
Test the comment regression with two repeatedly observed comments, four initial comments and restart. Exercise cancellation before/after acceptance, deferred-send failure, generation failure after accepted input, queue edits and unsupported steering. Specify load, context/call bounds and acceptable latency/critical-event thresholds before measurement. Track p95 useful-answer latency, action delay, critical-event misses, unnecessary wakes, oldest pending age and total model cost against foreground-only and deterministic-inbox baselines.
Also test a completed turn without its required artifact, a review queued before a crash, the same review arriving internally and through SCM, and foreground replacement while the previous background process remains alive. Idle background/foreground activity must produce no triage loop. Three routine updates plus one human request must dispatch one bounded-context turn. Exercise steer refusal, successful steer with receipt-write failure, every queue/outbox drain, and unexpected background approval. An interior deletion causes one reconciliation; an empty log follows epoch/watermark state, not emptiness alone.
Remaining scope and limits
Defer multi-host failover, multiple background replicas, arbitrary third-party automation tools, full transcript merging, automatic TUI foreground context insertion and a general workflow engine. SQLite, one background supervisor per orchestrator and existing workers are sufficient for this release.
Model urgency mistakes, unobserved history, irreversible accepted effects and remote exactly-once execution are limits, not deferred features. Conflicting actions outside AO's managed path remain possible. Shared provider failure can affect both sessions. This architecture does not itself improve coding accuracy or CI speed.
Keep a stable prompt prefix and append bounded deltas where possible. Checkpoint/rebuild when context grows. Cache hits and net savings need measurement; two sessions can cost more. The architecture earns its place if concurrent interaction and authorized progress improve without unacceptable misses, duplicate actions or context cost.
All reactions