Skip to content

bide v0.10.0

Choose a tag to compare

@github-actions github-actions released this 03 Oct 02:32
· 25 commits to main since this release
e417983

Runs you can steer, cancel and recover exactly as they started.

v0.10.0 completes the engine half of the pre-1.0 redesign. A run takes a Message and per-run options, journals them when it starts, and every later drive (a recovery drive included) runs under them. Runs can be cancelled and inspected from the journal, tools are described by one specification, construction is checked when the agent is built, and recovery hands each run to the code that knows how to drive it. Four new TLA+ models check the tool-call state machine, the run lifecycle, delegation and sessions, and P14's contract is written rule by rule against them.

  • One run entry point, with options that survive recovery. RunMessage, ResumeRun, StreamMessage (with AgentStream.Result()), RunTypedMessage[T] (a Go 1.27 generic method) and Session.SendMessage/SendMessageOnce take a Message input, so a run can start from text and images, and return a Result on every outcome once the run ID is valid: an answer, a pause, a halt, a failure, a saga's abort or a cancellation. The run's input and options (turn limit, token budget, system prompt, sampling, tool choice, saga, tool filter, output mode, typed schema, principal) are journaled in run:start, and every later drive runs under them. A later drive may raise or lower a limit, journaled as an amendment; any other different setting is ErrConfig before any model call. These are the transitional names of the 1.0 API, which the rewrite renames Run, Resume, Stream, RunTyped, Send and SendOnce.
  • Cancel and Status. agent.Cancel(ctx, j, runID, reason) stops a run at its next check: when a drive starts, at every turn boundary and after each won side-effect claim, so no side effect fires under a claim won after the cancellation. A cancelled saga rolls back first and then ends run:cancelled. A sub-agent's run, a programmatic sub-run and a delegation read their tree root's cancellation, so cancelling the root stops the whole tree, and plan flows honour Cancel too. The first end marker in journal order is a run's end for every reader. agent.Status reports a run's state (RunState) from one Load.
  • A per-run tool filter, enforced at dispatch. WithToolFilter(names...) restricts a run to some of the agent's tools. It is journaled with the run and checked when each call is dispatched, so a call outside it records an error result and never runs, on the first drive and on every resume. WithToolChoice(ToolChoice{Mode: "none"}) is enforced the same way.
  • Recovery that knows what it is resuming. Recover and RecoverLoop read each run's run:start under its lease and pass it to an agent.Resumer. ResumeAgent(a) drives an agent's plain runs and sagas, ResumeTyped[T](a) its typed runs, and ResumeAny combines them with your own. A run that has not started yet, or that no Resumer drives, is reported once per process and read again on every pass. RecoverLoop runs a second loop that takes over only runs whose lease has lapsed, so a dead holder's run is picked up within about one interval however many halted runs the store holds, and Leaser.ReapLeases clears the leases no pass will take over. A recovery pass is cheaper than before although it reads every driven run's start.
  • Tools described by one specification. agent.ToolSpec holds everything the agent knows about a tool (name, title, description, input and output schemas, safety, approval gate, timeout), read once when the tool is registered. Tool options set it: WithApproval, WithTimeout, WithTitle, WithOutputSchema and WithSafety. Safety is plain, comparable data (ReadOnly, Idempotent), the approval gate has its own field, with agent.SingleApproval() for one decision and ApprovalPolicy for m-of-n, and both are journaled with each tool result, so a saga rollback reads what the call ran under. Tool middleware receives an agent.ToolCall (Use, Spec, RunID). MCP tools gain mcp.WithApproval, their title, their output schema and their timeout.
  • Tool timeouts and tool calls that never misreport. WithTimeout bounds a call, middleware included. A late result is recorded; a late error has an unknown outcome, so a side effect halts on resume rather than record a failure for an effect that may have landed. Each call's state is tracked by compare-and-swap: a side effect whose tool began and did not itself fail is never recorded as a known failure, a result needs positive proof that no invocation is still running, and no invocation of next begins the tool after the chain has returned. ErrToolNotCalled marks a call known never to have reached its tool, and SagaAborted.UnknownOutcome lists saga steps that may have committed.
  • Construction checked when the agent is built. agent.Build(model, journal, opts...) and Agent.With(opts...) return every configuration problem as ErrConfig up front: nil parts, duplicate or reserved tool names, a tool name the model's provider refuses (ToolRules, declared by the OpenAI, Anthropic and Gemini adapters), non-object schemas, invalid approval policies and negative limits. Option scopes are interfaces, so an option passed where it does not apply does not compile. Decorators that hide a tool's approval gate or timeout are refused unless they implement Spec or Unwrap.
  • Programmatic sub-runs that a saga's rollback reaches. agent.RunInfoFrom(ctx) describes the tool call a context belongs to, and RunInfo.SubRunFor(name) names a sub-run started from that call. With WithSubRuns, a saga links each such sub-run in its journal before the sub-run writes anything, and its rollback compensates the sub-run's writes with the declared agent's compensators, latest first.
  • Retrieval as a journaled engine step. WithRetrieval(r, k) is an agent option: the agent retrieves for the run's user message once, journals the documents, and every model middleware sees the request with them in it. WithRetrievalRetry retries a failed retrieval with backoff and full jitter, recording only the attempt that succeeded.
  • Delegation that rolls back under its own authority. A saga whose delegations were minted under several grants (a root grant rotated or expired between drives) rolls back with all of them bound (audit.WithRollbackGrants), each child grant verified against its own parent. A delegation's rollback keeps the grant's expiry guard, and the rollback walks a sub-agent's run whatever the call's result.
  • Sessions under a lease. Over a store with leases, one worker drives a session turn at a time, so a turn under WithTokenBudget spends its budget once. Another worker returns the turn's recorded answer, or ErrTurnContended while it runs. A stale handle reads the journal before it refuses a new message, two callers on one handle record a turn once, and a cancelled turn is closed by the next message's Send.
  • Four new TLA+ models, with the code kept in step. Model 9 checks the tool-call state machine, model 10 the run lifecycle and recovery, model 11 delegation, sub-run authority and saga trees, and model 12 sessions, all under store faults, crashes and cancellations. Between them they found T1 to T6, L1 to L7, D1 to D3 and S1 to S4, and every one is fixed in this release and kept as a regression configuration. Model 10 was extended to state P14's rules before any of P14's code was written, and P14 implements a 16-rule contract from models 10 and 12, each rule with a test committed failing before its code. Code the models describe carries region markers, and CI (modelsync) fails a change to marked code that leaves its model unchanged without saying why. CI runs the models and the Linux Go tests in parallel shards with no check dropped, so every model is checked in about 4 minutes, and the larger configurations run nightly. Formal verification gives an overview of every model, what it guarantees and every bug it caught.
  • Claim safety proved beyond the model checker's bounds. The claim model has an inductive invariant that Apalache checks nightly, beside TLC. For two drivers over two processes it proves AtMostOnce and NotStartedExclusive on one call (attempts 0..3, 8 claim ids), and all four claim properties (AtMostOnce, NotStartedExclusive, NoLiveOverride and AtMostOncePerIntent) with halt resolution and the caller's second call (attempts 0..3, 6 claim ids), at any depth and for any number and mix of faults within those claim ids and attempts. It covers the claim protocol without the approval gate, and under the lease check it assumes no plain run holds the live attempt at the check (PlainRunIdleAtCheck). See the Apalache results.
  • Every governance example's gsm machine checked by the proof. The required gsm machine gate runs the two checkers extracted from gsm's Coq/Rocq proof, one on a machine's step tables and one on its rules, on every gsm machine the governance examples build, and the examples declare their rules with gsm's combinators so both checkers verify each one. It checks the examples' machines in CI, not the machines an application builds. The gate's first run found that the mesh example's signals did not commute; a signal now only raises its line's level, and a test checks every pair of signals in both orders.
  • gsm v0.12.0, with the proof's oracles in every Build. bide's governance tier requires gsm v0.12.0. Its Build checks compensation commutativity exactly for every pair it covers (every event pair, or only the pairs declared with Independent, in which case the guarantee covers runs whose events are reordered only across those pairs), with no footprint shortcut, and returns a machine only after the table oracle, generated from gsm's Coq/Rocq proof, re-checks it in-process; for combinator rules inside its fragment and within a cost cap, the rules oracle re-checks it from the rules as well. A federation's own conditions are checked by gsm's Go code. govern.CertifyConvergence reports Converges only for a machine Build returned. A convergence verdict or certificate recorded under gsm v0.11.0 is not covered: rebuild the machine and record a new one (what the checks cover).
  • Smaller records, cheaper recovery. agent.Record is 256 bytes, down from 368, which moves it into a smaller Go size class: rarely set fields (a model turn's metadata, an approver's signature) sit behind pointers, and the journal encoding is byte for byte the same. A recovery pass reads the commonest run:start without the JSON decoder, and a leased drive starts its renewal goroutine only when its first renewal is due.
  • Performance. On a standard 4-vCPU GitHub runner (AMD EPYC 7763, median of 21, run 37084060735), the overhead scenario runs ~24,700 runs/s at a mean run latency of 10.4 ms (p90 19.9 ms, p99 42.1 ms), and 20,000 runs with 5,000 in flight and 50 ms model calls finish in about 0.99 s (~20,100 runs/s, mean 248 ms, p90 309 ms, p99 414 ms). Against v0.9.0 on one machine (bench.yml A/B, AMD EPYC 7763, 21 interleaved runs per binary, run 37084062697), overhead throughput is flat (-0.5%) with p90 and p99 down 17%, and fan-out throughput is up 4.7%, with p90 down 7% and p99 down 9% (452 to 410 ms). In the Go benchmarks (benchstat, 10 runs per ref, same job), a recovery pass over 10,000 runs takes 32% less time with 11% fewer allocations; a model turn takes 12.7% more time and a side-effect tool call 7.2% more, with 3 to 5% fewer bytes per op, as the release adds store reads per turn and per side-effect call for cancellation (the counting-store budget in the CHANGELOG). On a laptop (Apple M1 Pro, 10 cores, a shared machine under load, 11 interleaved runs per binary), overhead runs ~52,200 runs/s (mean 4.9 ms, p90 9.7 ms, p99 18.9 ms) and fan-out ~38,500 runs/s (mean 130 ms, p90 136 ms, p99 164 ms), within 3.3% of v0.9.0's throughput on both. Latency is reported as the mean, p90 and p99, since the closed-loop harness's p50 is bimodal (cmd/bench).

API changes:

  • agent.Recover and agent.RecoverLoop take an agent.Resumer, func(ctx, runID string, start RunStart) error; use agent.ResumeAgent(a).
  • RunStart.Input is a Message; Sampling and ToolChoice marshal with snake_case JSON names.
  • Typed runs journal their typed start, so a typed run resumes through RunTypedMessage or ResumeTyped, not Run.
  • A run's first end marker in journal order is its end; ErrNotStarted wraps no category.
  • A finished run returns its recorded answer only to a drive with the input its run:start recorded.
  • Safety is {ReadOnly, Idempotent}: the approval gate is WithApproval(SingleApproval()) or WithApproval(&ApprovalPolicy{...}), and Safety.IdempotencyKey, RequiresApproval and Approval are removed. ApprovalPolicy literals need field names, and one with no approvers is ErrConfig.
  • ToolHandler is func(ctx, ToolCall) (json.RawMessage, error); middleware passes next the call it was given. ToolSafety, WithToolSafety and ToolErrorText are removed.
  • Request.Tools is []ToolSpec and ToolsDigest takes []ToolSpec; tool-result records carry Safety and Approval instead of ReadOnly.
  • govern.EventTool, govern.FederatedEventTool and audit.AttenuatingSubAgent take config structs; govern.AttestedEventTool is removed.
  • mcp.WithSafety sets only ReadOnly and Idempotent (gate with mcp.WithApproval), and mcp.Tools refuses a non-object outputSchema.
  • Tool decorators that hide an approval gate or timeout, wrappers that are also Compensators, and sub-agent wrappers with their own Timeout or Safety are ErrConfig; invalid tool options panic with ErrConfig.
  • SagaAborted.UnknownOutcome lists saga steps that may have committed; a side effect that began is never recorded as a known failure.
  • RunFilter.Admits takes a lapsed function, and Leaser has ReapLeases.
  • Record's model-turn fields (Finish, RawFinish, Model, PromptDigest, ToolsDigest) and approver fields (Approver, ApproverAlg, Signature) move behind *ModelTurn and *ApproverSignature; read them with the methods of the same names.
  • WithRetrieval is an agent option, RetrievalTool takes a name and description, and trace.Instrument returns an agent option.
  • The context decorators are ContextWithIdentity, ContextWithWaker and ContextWithClock; the With names are agent and run options. WithNow is WithClock and StepSafety is WithSafety.
  • Parallel takes its tasks as a slice and its cap as WithMaxConcurrency(n); Lease, Recover and RecoverLoop take their own option types.
  • govern requires gsm v0.12.0; govern.CertifyConvergence reports Converges only for a machine Build returned.
  • RunScope and InSaga are replaced by RunInfoFrom; a plain run started from a saga's tool call is not a saga.

The CHANGELOG has the details of each change. The deprecated entry points (Agent.Run with a string, RunSaga, Stream, RunTyped, Session.Send and SendOnce, agent.New and its builder methods) still work in v0.10.0; the 1.0 rewrite removes them.

govern requires gsm v0.12.0: rebuild your machines under it, since its Build may refuse one v0.11.0 accepted; gsm's v0.12.0 release notes list the other changes.

Upgrading journals across the pre-release

Before 1.0, a run is not promised to resume across releases (the journal format tag stays bide.journal.v1-dev), so finish or resolve runs before upgrading. When runs are in flight across the upgrade from v0.9.0:

  • Recovery does not guess between a plain and a typed run started by v0.9.0: ResumeAgent and ResumeTyped report such a run as ErrNotResumable. Recover typed v0.9.0 runs with a Resumer of your own that knows which entry point started each run.
  • Tool results journaled by v0.9.0 carry no safety, so a saga rollback over them treats every completed call as a write.
  • A saga that holds an ungranted audit.AttenuatingSubAgent delegation cannot be rolled back after the upgrade (the rollback stops with ErrProtocol): finish or roll back such sagas first.
  • A convergence verdict or govern.ConfluenceCertificate recorded under gsm v0.11.0 is not covered by v0.12.0's fix: rebuild the machine and record a new certificate.
  • On Postgres, Open creates two indexes on the leases table on first open; on a store whose role does not own the tables, the owner opens it once to create them.

Install:

go get github.com/bide-ai/bide@v0.10.0

Library modules, at the same version:

go get github.com/bide-ai/bide/store/sqlite@v0.10.0
go get github.com/bide-ai/bide/store/postgres@v0.10.0
go get github.com/bide-ai/bide/mcp@v0.10.0
go get github.com/bide-ai/bide/trace@v0.10.0
go get github.com/bide-ai/bide/codec/gcf@v0.10.0
go get github.com/bide-ai/bide/govern@v0.10.0
go get github.com/bide-ai/bide/govern/sqlitelog@v0.10.0
go get github.com/bide-ai/bide/govern/redislog@v0.10.0
go get github.com/bide-ai/bide/govern/postgreslog@v0.10.0

Docs: https://bide-ai.com. Formal models: spec/tla and the formal verification overview.