Skip to content

Releases: bide-ai/bide

bide v0.12.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 03:32
a4c1e48

Federations checked for event order.

v0.12.0 moves bide's governance tier to gsm v0.13.0.

  • Event order across registries is checked. Federation.Build now rejects a federation where the order of events across registries could change the outcome (the C1 and C2 checks), on trees, acyclic networks and monotone cycles. The federated convergence statement is machine-checked in this corrected form in normalization-confluence.
  • Certificates record delivery obligations. govern.ConfluenceCertificate carries what gsm reports a deployment must provide: pairs that need causal delivery, events that need deduplication, and writes clamped into a variable's range. bide's shared event log and ApplyOnce already provide the first two.

API changes:

  • bide requires gsm v0.13.0.
  • govern.ConfluenceCertificate gains PairsUndeclared, CausalOrderRequired, NotIdempotent and Saturations (types govern.EventPair, govern.Saturation); verify its certificates with bide-audit v0.12.0.
  • From gsm: Machine.Apply normalizes an invalid input before applying the event, and Embed no longer copies AllowMonotoneCycles.

The CHANGELOG has the details.

Install:

go get github.com/bide-ai/bide/agent@v0.12.0

Library modules, at the same version:

go get github.com/bide-ai/bide/store/sqlite@v0.12.0
go get github.com/bide-ai/bide/store/postgres@v0.12.0
go get github.com/bide-ai/bide/mcptools@v0.12.0
go get github.com/bide-ai/bide/trace@v0.12.0
go get github.com/bide-ai/bide/codec/gcf@v0.12.0
go get github.com/bide-ai/bide/govern@v0.12.0
go get github.com/bide-ai/bide/govern/sqlitelog@v0.12.0
go get github.com/bide-ai/bide/govern/redislog@v0.12.0
go get github.com/bide-ai/bide/govern/postgreslog@v0.12.0

Docs: https://bide-ai.com. Formal models: spec/tla and the formal verification overview.

bide v0.11.1

Choose a tag to compare

@github-actions github-actions released this 04 Oct 03:15
0d91de9

A smoother first run.

v0.11.1 is a small follow-up to v0.11.0, from a first-run test of the getting started guide.

  • A clear error for a missing API key. With model/openai and an empty key, a provider that needs one refuses the request; the error now says no API key was set, instead of reading like a bad key. A server that needs no key, such as Ollama, works as before.
  • Model errors say (model) once. A failed model call's error no longer repeats the suffix.
  • Getting started works without a key. The guide runs the first agent offline with agenttest's scripted model, installs with go get github.com/bide-ai/bide/agent@latest, runs the core examples with go run github.com/bide-ai/bide/examples/<name>@latest, and says what a run prints and how a crashed run continues.

No API changes. The CHANGELOG has the details.

Install:

go get github.com/bide-ai/bide/agent@v0.11.1

Library modules, at the same version:

go get github.com/bide-ai/bide/store/sqlite@v0.11.1
go get github.com/bide-ai/bide/store/postgres@v0.11.1
go get github.com/bide-ai/bide/mcptools@v0.11.1
go get github.com/bide-ai/bide/trace@v0.11.1
go get github.com/bide-ai/bide/codec/gcf@v0.11.1
go get github.com/bide-ai/bide/govern@v0.11.1
go get github.com/bide-ai/bide/govern/sqlitelog@v0.11.1
go get github.com/bide-ai/bide/govern/redislog@v0.11.1
go get github.com/bide-ai/bide/govern/postgreslog@v0.11.1

Docs: https://bide-ai.com.

bide v0.11.0

Choose a tag to compare

@github-actions github-actions released this 04 Oct 01:54
3135a0f

The API under its new names.

v0.11.0 completes the API redesign's rename: the transitional names of v0.10.0 become the API, and the old API they stood beside is gone. One way to build an agent, one way to run it, one journal type, one tool interface. The protocols the formal models check are unchanged; what changes is the surface, a few refusals that turn silent mistakes into errors when the agent is built, and audit anchoring after every write.

  • One way to build and run. agent.New(model, journal, opts...) returns (*Agent, error) and checks every configuration problem up front. Run, Stream, Resume, RunTyped and Session.Send/SendOnce take a Message and run options and return a *Result; a saga is Run(..., WithSaga()).
  • One journal type. Every function that took a Durable takes a *agent.Journal (agent.NewJournal(store)), and the journal verbs are its methods: j.Step, j.Parallel, j.Signal, j.Enqueue, j.AnswerInterrupt. A store wrapper wraps the Store, and implements Unwrap() Store when it passes keys through.
  • One tool interface. A tool is Spec() ToolSpec and Call. A decorator that embeds a tool inherits its spec, approval gate and timeout included, and overrides in its own Spec.
  • Mistakes refused when the agent is built. A tool with a Safety(), Name(), Description() or ArgsSchema() method that disagrees with its Spec (an override written for the old interface, which nothing reads any more) is ErrConfig from New and from a plan flow, so a decorator's own side effect is never re-run on resume.
  • Audit anchoring after every write. audit.AuditedStore wraps a Store and anchors after every insert, a write that committed while reporting an error included; a write it cannot anchor (its context ended first, or the publish failed) is reported to OnError when one is set, and AuditedStore.Reanchor anchors that run later.

API changes:

  • Build is New(model, journal, opts...) (*Agent, error); the old New(model, store, tools...) and the builder methods (Use, UseTool, WithMaxTurns, SetMaxConcurrency, ...) are removed: pass the options to New, or derive a copy with With.
  • RunMessage, StreamMessage, ResumeRun, RunTypedMessage and Session.SendMessage/SendMessageOnce are Run, Stream, Resume, RunTyped and Session.Send/SendOnce. The string entry points (Run and Stream on a string, RunSaga, RunResult, RunSagaResult, StreamSaga, RunTypedNative, AgentStream.Final) are removed. AgentStream and AgentEvent are RunStream and RunEvent; audit.Record is audit.RecordStream.
  • The Durable interface and the stores' Do, History and Journal shims are removed; use *Journal. Step, Parallel, Signal, Enqueue (the channel Send) and AnswerInterrupt (Resume[T]) are *Journal methods.
  • PendingApproval, Interrupted, Awaiting, Sleeping and ResumeHalt are ApprovalPending, InterruptPending, SignalPending, TimerPending and OutcomeUnknown; ResolveHaltRef is ResolveHalt; ResolveStepHalt and ApproveAs are removed.
  • ContextWithIdentity, ContextWithWaker and ContextWithClock are removed: pass WithIdentity, WithWaker or WithClock as a run option. WithIdentity refuses an empty identity.
  • The Tool interface is Spec() and Call; Name, Description, ArgsSchema, Safety and SpecOf are gone. Func takes its safety as WithSafety; Func, CompensatedFunc, SubAgent and RetrievalTool return (Tool, error), with MustFunc, MustCompensatedFunc, MustSubAgent and MustRetrievalTool that panic instead.
  • New refuses a tool whose old-interface method disagrees with its Spec; plan flows refuse it too.
  • Package mcp is mcptools (github.com/bide-ai/bide/mcptools).
  • The scripted model and its turns are in agent/agenttest, with new helpers: MemJournal, MustJournal, MustNew, Must, Answer and IdentityContext.
  • plan's Register* functions are methods of *plan.Registry.
  • Run on a run journaled as a saga drives it as the saga, without WithSaga.
  • audit.AuditedStore wraps a Store, anchors after every insert but the header's, and gains Reanchor(ctx, runID); cmd/bide-audit refuses an export that does not start with a journal header.

The CHANGELOG has the details of each change.

Install:

go get github.com/bide-ai/bide@v0.11.0

Library modules, at the same version:

go get github.com/bide-ai/bide/store/sqlite@v0.11.0
go get github.com/bide-ai/bide/store/postgres@v0.11.0
go get github.com/bide-ai/bide/mcptools@v0.11.0
go get github.com/bide-ai/bide/trace@v0.11.0
go get github.com/bide-ai/bide/codec/gcf@v0.11.0
go get github.com/bide-ai/bide/govern@v0.11.0
go get github.com/bide-ai/bide/govern/sqlitelog@v0.11.0
go get github.com/bide-ai/bide/govern/redislog@v0.11.0
go get github.com/bide-ai/bide/govern/postgreslog@v0.11.0

Docs: https://bide-ai.com. Formal models: spec/tla and the formal verification overview.

bide v0.10.0

Choose a tag to compare

@github-actions github-actions released this 03 Oct 02:32
e417983

Runs you can steer, cancel and recover exactly as they started.

v0.10.0 completes the engine half of the pre-1.0 redesign. A run takes a Message and per-run options, journals them when it starts, and every later drive (a recovery drive included) runs under them. Runs can be cancelled and inspected from the journal, tools are described by one specification, construction is checked when the agent is built, and recovery hands each run to the code that knows how to drive it. Four new TLA+ models check the tool-call state machine, the run lifecycle, delegation and sessions, and P14's contract is written rule by rule against them.

  • One run entry point, with options that survive recovery. RunMessage, ResumeRun, StreamMessage (with AgentStream.Result()), RunTypedMessage[T] (a Go 1.27 generic method) and Session.SendMessage/SendMessageOnce take a Message input, so a run can start from text and images, and return a Result on every outcome once the run ID is valid: an answer, a pause, a halt, a failure, a saga's abort or a cancellation. The run's input and options (turn limit, token budget, system prompt, sampling, tool choice, saga, tool filter, output mode, typed schema, principal) are journaled in run:start, and every later drive runs under them. A later drive may raise or lower a limit, journaled as an amendment; any other different setting is ErrConfig before any model call. These are the transitional names of the 1.0 API, which the rewrite renames Run, Resume, Stream, RunTyped, Send and SendOnce.
  • Cancel and Status. agent.Cancel(ctx, j, runID, reason) stops a run at its next check: when a drive starts, at every turn boundary and after each won side-effect claim, so no side effect fires under a claim won after the cancellation. A cancelled saga rolls back first and then ends run:cancelled. A sub-agent's run, a programmatic sub-run and a delegation read their tree root's cancellation, so cancelling the root stops the whole tree, and plan flows honour Cancel too. The first end marker in journal order is a run's end for every reader. agent.Status reports a run's state (RunState) from one Load.
  • A per-run tool filter, enforced at dispatch. WithToolFilter(names...) restricts a run to some of the agent's tools. It is journaled with the run and checked when each call is dispatched, so a call outside it records an error result and never runs, on the first drive and on every resume. WithToolChoice(ToolChoice{Mode: "none"}) is enforced the same way.
  • Recovery that knows what it is resuming. Recover and RecoverLoop read each run's run:start under its lease and pass it to an agent.Resumer. ResumeAgent(a) drives an agent's plain runs and sagas, ResumeTyped[T](a) its typed runs, and ResumeAny combines them with your own. A run that has not started yet, or that no Resumer drives, is reported once per process and read again on every pass. RecoverLoop runs a second loop that takes over only runs whose lease has lapsed, so a dead holder's run is picked up within about one interval however many halted runs the store holds, and Leaser.ReapLeases clears the leases no pass will take over. A recovery pass is cheaper than before although it reads every driven run's start.
  • Tools described by one specification. agent.ToolSpec holds everything the agent knows about a tool (name, title, description, input and output schemas, safety, approval gate, timeout), read once when the tool is registered. Tool options set it: WithApproval, WithTimeout, WithTitle, WithOutputSchema and WithSafety. Safety is plain, comparable data (ReadOnly, Idempotent), the approval gate has its own field, with agent.SingleApproval() for one decision and ApprovalPolicy for m-of-n, and both are journaled with each tool result, so a saga rollback reads what the call ran under. Tool middleware receives an agent.ToolCall (Use, Spec, RunID). MCP tools gain mcp.WithApproval, their title, their output schema and their timeout.
  • Tool timeouts and tool calls that never misreport. WithTimeout bounds a call, middleware included. A late result is recorded; a late error has an unknown outcome, so a side effect halts on resume rather than record a failure for an effect that may have landed. Each call's state is tracked by compare-and-swap: a side effect whose tool began and did not itself fail is never recorded as a known failure, a result needs positive proof that no invocation is still running, and no invocation of next begins the tool after the chain has returned. ErrToolNotCalled marks a call known never to have reached its tool, and SagaAborted.UnknownOutcome lists saga steps that may have committed.
  • Construction checked when the agent is built. agent.Build(model, journal, opts...) and Agent.With(opts...) return every configuration problem as ErrConfig up front: nil parts, duplicate or reserved tool names, a tool name the model's provider refuses (ToolRules, declared by the OpenAI, Anthropic and Gemini adapters), non-object schemas, invalid approval policies and negative limits. Option scopes are interfaces, so an option passed where it does not apply does not compile. Decorators that hide a tool's approval gate or timeout are refused unless they implement Spec or Unwrap.
  • Programmatic sub-runs that a saga's rollback reaches. agent.RunInfoFrom(ctx) describes the tool call a context belongs to, and RunInfo.SubRunFor(name) names a sub-run started from that call. With WithSubRuns, a saga links each such sub-run in its journal before the sub-run writes anything, and its rollback compensates the sub-run's writes with the declared agent's compensators, latest first.
  • Retrieval as a journaled engine step. WithRetrieval(r, k) is an agent option: the agent retrieves for the run's user message once, journals the documents, and every model middleware sees the request with them in it. WithRetrievalRetry retries a failed retrieval with backoff and full jitter, recording only the attempt that succeeded.
  • Delegation that rolls back under its own authority. A saga whose delegations were minted under several grants (a root grant rotated or expired between drives) rolls back with all of them bound (audit.WithRollbackGrants), each child grant verified against its own parent. A delegation's rollback keeps the grant's expiry guard, and the rollback walks a sub-agent's run whatever the call's result.
  • Sessions under a lease. Over a store with leases, one worker drives a session turn at a time, so a turn under WithTokenBudget spends its budget once. Another worker returns the turn's recorded answer, or ErrTurnContended while it runs. A stale handle reads the journal before it refuses a new message, two callers on one handle record a turn once, and a cancelled turn is closed by the next message's Send.
  • Four new TLA+ models, with the code kept in step. Model 9 checks the tool-call state machine, model 10 the run lifecycle and recovery, model 11 delegation, sub-run authority and saga trees, and model 12 sessions, all under store faults, crashes and cancellations. Between them they found T1 to T6, L1 to L7, D1 to D3 and S1 to S4, and every one is fixed in this release and kept as a regression configuration. Model 10 was extended to state P14's rules before any of P14's code was written, and P14 implements a 16-rule contract from models 10 and 12, each rule with a test committed failing before its code. Code the models describe carries region markers, and CI (modelsync) fails a change to marked code that leaves its model unchanged without saying why. CI runs the models and the Linux Go tests in parallel shards with no check dropped, so every model is checked in about 4 minutes, and the larger configurations run nightly. Formal verification gives an overview of every model, what it guarantees and every bug it caught.
  • Claim safety proved beyond the model checker's bounds. The claim model has an inductive invariant that Apalache checks nightly, beside TLC. For two drivers over two processes it proves AtMostOnce and NotStartedExclusive on one call (attempts 0..3, 8 claim ids), and all four claim properties (AtMostOnce, NotStartedExclusive, NoLiveOverride and AtMostOncePerIntent) with halt resolution and the caller's second call (attempts 0..3, 6 claim ids), at any depth and for any number and mix of faults within those claim ids and attempts. It covers the claim protocol without the approval gate, and under the lease check it assumes no plain run holds the live attempt at the check (PlainRunIdleAtCheck). See the Apalache results.
  • Every governance example's gsm machine checked by the proof. The required gsm machine gate runs the two checkers extracted from gsm's Coq/Rocq proof, one on a machine's step tables and one on its rules, on every gsm machine the governance examples build, and the examples declare their rules with gsm's combinators so both checkers verify each one. It checks the examples' machines in CI, not the machines an application builds. The gate's first run found that the mesh example's signals did not commute; a signal now only raises its line's level, and a test checks every pair of signals in both orders.
  • gsm v0.12.0, with the proof's oracles in every Build. bide's governance tier requires gsm v0.12.0. Its Build checks compensation commutativity exactly for every pair it covers (every event pair, or only the pairs declared with Independent, in which case the guarantee covers runs whose events are reordered only across those pairs), with no footprint shortcut, and returns a machine only after the table oracle, generated from gsm's Coq/Rocq proof, re-checks it in-process; for combinator rules ins...
Read more

bide v0.9.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 02:43
4faaaf3

Correction (2026-10-02). This release's README overstated gsm's convergence guarantee: gsm v0.11.0 could certify some machines as convergent when they aren't (when one event reads a variable another event writes), and the proof-extracted checkers didn't run in Build or CI. The proof is correct; the implementation skipped a precondition. Fixed in gsm#2; details in bide#142.

A claim protocol checked by a model checker, on a core built for 1.0.

v0.9.0 is the release where bide's coordination protocols are checked, not only tested. TLA+ models of claims, approvals, flows, spend and the bide protocol run with the TLC model checker on every pull request, and the engine underneath them is rebuilt around a storage port and a journal with its own format. Model calls are typed values with exact spend accounting, flows run on durable Steps, proofs commit to the bytes the journal stores, and every audit entry point works under any signature scheme.

  • The claim protocol, model-checked in CI. TLA+ models (written in PlusCal) cover attempt claims, not-started records, numbered retries, the resume gate and halt resolution (model 1); the approval gate with 1-of-1 and m-of-n tallies and approvers' key sets (model 1b); flow semantics (model 7); spend accounting (model 8); and the bide protocol's claim rules for remote tool calls (model 2). TLC explores every interleaving within each configuration's bounds, and Models is a required check on every pull request. Each counterexample the checker finds becomes a deterministic Go regression test, the Go fault-schedule exploration of concurrent claims is deterministic too (with a nightly full-bound run), and the model's findings are now properties the engine holds: each effect fires at most once, at most one attempt of a call is live, a resolver that never overrides a live driver, and a claim's winner that never halts on a loser's behalf. See spec/tla.
  • A Store/Journal core. agent.Store is a small storage port (Insert, Get, Load) with its requirements numbered A1 to A8 on the type, and agent.Journal owns everything above it: memoization, the record encoding and salt, attempt claims, not-started records and recording an outcome after the caller's context ends. Every journal opens with an @journal header naming its format, so a run in another format is refused before anything is read or written. agent/storetest checks any store against every requirement, RunFilter lets SQL stores list only unfinished runs, and a counting store holds the engine to an exact budget of store round trips per operation. Journals in one process share in-flight steps, and a claim whose write may not have committed is recorded as not started, so a re-drive re-attempts the effect rather than halting over one that never ran.
  • A sealed pause contract. Every pause satisfies agent.Pause, with five kinds (ApprovalPending, InterruptPending, SignalPending, TimerPending, OutcomeUnknown) and verbs named for the pause they answer (SubmitDecision, AnswerInterrupt, Enqueue). agent.ResolveHaltRef resolves tool and Step halts alike, knows whether a halt was crashed or contended, and refuses to resolve an effect a live driver may still be running.
  • Typed model calls and exact spend. A model call is an agent.ModelCall value and returns an agent.ModelResponse, so middleware passes the call on instead of threading engine state through the context, and agent.CallModel sends one outside an agent through the same checks. Hooks are append-only with exactly one After per Before, the spend meter sits beneath every middleware, and requests are numbered across retries and hedged targets. A run that ends waits briefly for requests still in flight and journals their usage as late spend, so hedge losers and abandoned requests count toward Result.Spend, the budget and Replay. Each turn journals its finish reason, the model that answered, and digests of the prompt and tool set it was sent. The stream sink is claimed by one request per turn, so a hedged turn streams live from its first target.
  • Flows on durable Steps, with conformance. Every plan node runs as an agent.Step, so flows get the Step's claim protocol, pause guard and halt resolution (plan.Flow.ResolveHalt). A flow run records what it is (RunKindFlow, its name and canonical input) and run:complete when it finishes, so recovery skips finished flows and a later drive returns the recorded output with one point read. Steps inside a loop body are scoped per iteration, and Conform replays a run's routing from its recorded choices. A fault-schedule exploration of the lowering (crashes, store errors and cancellations at every write, across processes) runs as a permanent test.
  • Proofs over the bytes the journal stores. A journal leaf is its tag followed by the record's stored bytes (Record.Raw), so a record written by a later release verifies unchanged, and audit.ExportJournal hands an auditor exactly those bytes. Every artifact names its format (proofs v3, evidence v5, STH v5), and audit.UnmarshalStrict checks it at every depth. Projections refuse a journal they cannot represent exactly, with ErrRedacted or ErrMalformed, and verifiers return an error that wraps ErrNotVerified, ErrFormat or ErrMalformed.
  • Signature agility and post-quantum hybrid. Every audit entry point takes an audit.Signer or audit.Verifier, so ed25519, ML-DSA-65 (FIPS 204) and the ed25519 plus ML-DSA-65 hybrid work throughout: tree heads, evidence, run certificates, approvals, grants, the standalone audit/verify package and bide-audit. The scheme is part of the signed bytes, keys have one text form (<alg>:<hex>), and each half of a hybrid signature signs its own label, so neither half verifies on its own. An m-of-n decision is journaled with the scheme it was signed under.
  • Approval key identity. An approver's verifier reports the identities of its signing keys (KeyIDs), and the m-of-n gate counts one seat per signing key: a policy whose approvers share a key is refused on every evaluation, by the gate, TallyApprovals, audit.VerifyApprovals and bide-audit verify-approvals alike. Ed25519 keys are checked for canonical encoding and full order (audit.CheckEd25519PublicKey), so every accepted key is exactly one key.
  • The bide protocol, designed for SDKs. bide.protocol.v1 is the accepted design for using bide from other languages: the Go engine stays the only writer of the journal, claims and proofs, while Python and TypeScript SDKs run tools and answer pauses, as long-lived workers or serverless functions. Its claim rules pass model 2. Implementation follows the engine hardening of the pre-1.0 redesign.
  • Postgres stores that hold nothing between round trips, in a schema you pin. Every store/postgres lease call and journal insert, and every govern/postgreslog append, is one statement that Postgres commits before it replies. An insert takes its position from a next_seq plpgsql function under a transaction-level advisory lock that ends with the statement, so a stalled client holds no lock and another node takes a run over one TTL after its last committed renewal. Every name the stores send is qualified: tables and next_seq with the store's schema, and every function, type and operator with pg_catalog (operators as OPERATOR(pg_catalog.<op>)), and a static check holds every statement to that rule. WithSchema pins the schema so the search path plays no part, and pinning it is the recommended setup. Open checks that the tables carry the unique indexes the statements rely on, and that next_seq has the expected definition, owner and search_path.
  • Recovery that drives only unfinished runs. Holding a run's lease, Recover and RecoverLoop check its end-of-run markers again before calling resume, so a run finished by another lease holder after the pass listed it is left alone, for three point reads per driven run.
  • Performance. On a standard 4-vCPU GitHub runner (AMD EPYC 7763, median of 21, run 36802470752), the overhead scenario runs ~25,400 runs/s at a mean run latency of 10.1 ms (p90 22 ms, p99 50 ms), and 20,000 runs with 5,000 in flight and 50 ms model calls finish in about 1.01 s (~19,900 runs/s, mean 251 ms, p90 321 ms, p99 447 ms). A same-ref run on the same CPU model put the noise within 2.7% on every one of these metrics. Throughput is flat to slightly up against v0.8.0: the Store/Journal core (#92) gained about 10%, and typed model calls (#104) gave back 3 to 8% as Record grew. The fan-out p99 is higher than v0.8.0's (447 vs 419 ms) while throughput and the mean improved: #92's sixth journal record per run (the @journal header) reshapes the latency distribution, and the Record shrink planned after P12 is expected to recover some of it. Latency is reported as the mean, p90 and p99, since the closed-loop harness's p50 is bimodal, sitting on a scheduling cliff (cmd/bench).

API changes:

  • Every journal starts with an @journal header (format bide.journal.v1-dev), so History returns it first and record indices shift by one; journals written by v0.8.0 and earlier are refused, not resumed.
  • store/sqlite names its tables bide_steps, bide_leases and bide_schema_version, and store/postgres its leases table bide_leases; stop every v0.8.0 node before starting v0.9.0.
  • agent.Lister.Runs takes a RunFilter and returns an iterator; agent.Capability takes a Store; Record.Salt and Record.Claim are read through Salt() and ClaimID().
  • The pause types are `Ap...
Read more

bide v0.8.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 07:06
629d47e

Runs that recover themselves, budgets that cover the whole tree.

v0.8.0 makes durable runs self-sustaining and fully accounted for. Recovery runs continuously, every run journals where it started, token budgets and spend reach every sub-agent and every request sent, and each library module now installs with go get.

  • Continuous recovery. agent.RecoverLoop keeps taking over the runs of a crashed holder with no further calls, with its own interval, concurrency and error handler. Leases are held under a per-call token, failed renewals retry well before the lease expires, and a drive that loses its lease stops with ErrLeaseLost. A multi-process test harness on Postgres kills, stalls and restarts workers to check it, and the Postgres stores run every write at read committed, whatever the deployment's default isolation.
  • Runs that know where they started. Each run journals its input and entry point (run:start), readable with agent.RecordedStart, so a Recover callback resumes a run with exactly what it began with. Resume decisions come from the journal: a recorded approval denial stays final, saga rollback compensates by the safety each call recorded, and a plan node re-runs only when it was retry-safe when attempted.
  • Sessions in their own namespace. A session's journal and turns live under "<id>>@session" and "<id>>@turn/<n>", so no run started with Run can share one, session ids may contain /, and agent.IsSessionRun identifies them. Timers inside sub-agents wake under names qualified by the sub-run, so sibling sub-agents can wait on timers of the same name.
  • One budget for the agent tree. WithTokenBudget on a parent covers its sub-agents, and Result.Usage and the new Result.Spend report the whole run, including retried, hedged and failed requests. middleware.CostMeter reports what was spent, a TurnRestarted event marks a retried streamed turn, and WithModelCallHook and WithModel add hooks and per-call routing around every request.
  • Precise model boundaries. Finish reasons are one typed, provider-neutral set (stop, tool_use, length, filtered) with the provider's own value in Finish.Raw, and a cut-off or filtered turn is reported as ErrOutputTruncated or ErrOutputFiltered. Responses are capped (WithMaxResponseBytes), adapters describe themselves through agent.ModelInfoOf, and the provider HTTP kit has its own package, model/provider, for writing adapters.
  • Typed tools, exactly as declared. Tool arguments decode strictly against their schema, schema.For describes precisely what encoding/json decodes, and a call whose outcome is lost mid-flight on a tool that is not retry-safe is recorded as unknown (ErrToolOutcomeUnknown). A call cancelled before its effect ran is re-attempted.
  • MCP with limits. mcp.WithSafety declares retry safety per tool, and WithCallTimeout, WithMaxResultBytes and WithMaxDescriptionBytes bound each server, with 1 MiB and 8 KiB defaults. Tool lists are checked for malformed and duplicate names, and modeltest.ToolNames checks provider name rules.
  • Audit artifacts for tooling. Proofs and evidence use versioned formats with snake_case JSON, signed tree heads carry checked timestamps (audit.CheckTimestamp, WithClockSkew), and bide-audit adds -json and -version with one exit-status scheme: 0 verified, 1 not verified, 2 usage, 3 no verdict, 4 input unusable.
  • Flows as versioned config. plan configs declare "version": 1 with snake_case keys, blocks can be named (Arm.Named, plan.BlockName), steps and joins take a context.Context, and a config safety can lower a node's retry safety, never raise it.
  • Evaluation you can gate on. A metric that cannot score a run leaves it unscored, eval.Compare returns a Comparison with a Gate and tolerances, and eval reports errors instead of panicking.
  • Retrieval across stores. RetrievalName and RetrievalDescription let one agent search several stores, and retrieved context is journaled once per run and kept in every model call, including after resume.
  • Journaling throughput at the v0.7.0 level. Recording a step decodes each record once and tags message parts without a re-encode, with the journal bytes unchanged. On a standard 4-vCPU GitHub runner (AMD EPYC 7763, median of 21), 20,000 runs with 5,000 in flight and 50 ms model calls finish in about 1.05 s (~19,000 runs/s), and the overhead scenario runs ~23,700 runs/s (cmd/bench).
  • Every module installable. govern is its own module, and the stores, mcp, trace, codec/gcf, govern and the governed-event logs are tagged with each release, so they install with go get at the same version as the core. CI checks the doc comment on every exported identifier and compiles every Go block in the README and docs against the current code.

API changes:

  • Journal keys encode the tool-use id (tool:<id>, sub-runs <parent>><id>); runs journaled by v0.7.0 do not resume.
  • ResolveHalt resolves tool calls only; use ResolveStepHalt for a Step.
  • Step names with a reserved engine prefix, run and session ids containing >, and plan step names containing : are ErrConfig, and Recover skips sub-runs.
  • Finish.Reason uses the neutral values stop, tool_use, length and filtered; a truncated or filtered turn fails with ErrOutputTruncated or ErrOutputFiltered, any other reason with ErrStreamProtocol.
  • A reply over the response cap fails with ErrResponseTooLarge, and a tool_use turn with no call fails.
  • Tool arguments for Func, SubAgent and final_answer decode strictly: missing required, unknown, case-variant, duplicate or trailing data is ErrToolArgs.
  • schema.For describes what encoding/json decodes and returns ErrUnsupportedType for kinds JSON cannot carry.
  • RunTyped[T] needs an object T, and RunTypedNative on Anthropic returns ErrConfig.
  • WithRetrieval sends context as a user message before the latest user turn, journaled once per run, and RetrievalTool and WithRetrieval require k >= 1.
  • WithTokenBudget on a parent covers its sub-agents, and Result.Usage and Result.Spend report the whole run.
  • agent.EmitMessage takes the turn's Usage.
  • plan step and join functions (Builder.Step, Join2, Join3, RegisterStep, RegisterJoin2, RegisterJoin3) take a context.Context.
  • Audit proof JSON is snake_case with a format field (bide.audit.proof.v2, bide.audit.evidence.v4), and other formats fail with audit.ErrFormat.
  • plan digests are bide.plan.topology.v2 (flows started under v1 do not resume), config JSON decodes strictly, and a config safety can lower retry safety but not raise it or clear approval.
  • plan configs need "version": 1 and snake_case keys, Topology JSON keys are snake_case, TopologyNode.Kind is TopologyNodeKind, and parse errors wrap ErrConfig.
  • Signed tree heads need a positive timestamp within the clock skew, empty evidence packages fail, and bide-audit input is capped at 256 MiB.
  • ApprovalPolicy.Validate refuses approver ids equal under Unicode normalization and case folding, and invalid UTF-8.
  • middleware.Retry and ToolRetry with n < 0 and Quorum with k above the voter count are ErrConfig, and EventLog.Events returns ErrProtocol on a gap.
  • eval.Judge passes only on an exact PASS, and a failed run never passes.
  • eval.RequiredRuns returns (int, error), and eval.Run returns an error for duplicate metric names.
  • eval.Matches takes a *regexp.Regexp, AgentRunner returns (RunFunc, error), Compare returns (Comparison, error), and GovernanceHeld predicates take a context.
  • The provider HTTP kit moves from agent to model/provider; WithToolResultCodec takes a provider.ToolResultCodec.
  • Finish.Reason is the typed agent.FinishReason.
  • eval Metric.Fn returns (bool, error), MetricStat reports Scored and Unscored, and reports are bide.eval.report.v2.
  • bide-audit exit statuses: 0 verified, 1 not verified, 2 usage, 3 no verdict, 4 input unreadable or unusable.
  • A session's journal and turn runs are "<id>>@session", "<id>>@turn/<n>" and "<id>>@event/<encoded key>"; sessions journaled by v0.7.0 open empty, and Recover skips them.
  • Every run journals run:start; resuming with a different input, a saga through Run (or a run through RunSaga), or SendOnce with a different input on an open turn is ErrConfig.
  • plan attempt markers record retry safety; a node re-runs on resume only if it was retry-safe then and now, and a marker written before v0.8.0 halts.
  • A recorded approval denial is final even if the tool's gate is later removed, loosened or made m-of-n.
  • Saga rollback treats a completed call as a write unless its result records it ran ReadOnly, and reports calls to unregistered tools as uncompensated.
  • plan.Retryable and the config safety "retryable" are removed; use plan.Idempotent and "idempotent".
  • govern is its own module: add github.com/bide-ai/bide/govern with go get.

Install:

go get github.com/bide-ai/bide@v0.8.0

Library modules, at the same version:

go get github.com/bide-ai/bide/store/sqlite@v0.8.0
go get github.com/bide-ai/bide/store/postgres@v0.8.0
go get github.com/bide-ai/bide/mcp@v0.8.0
go get github.com/bide-ai/bide/trace@v0.8.0
go get github.com/bide-ai/bide/codec/gcf@v0.8.0
go get github.com/bide-ai/bide/govern@v0.8.0
go get github.com/bide-ai/bide/govern/sqlitelog@v0.8.0
go get github.com/bide-ai/bide/govern/redislog@v0.8.0
go get github.com/bide-ai/bide/govern/postgreslog@v0.8.0

Docs: https://bide-ai.com. The approval guide is now Human approval (human-in-the-loop).

bide v0.7.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 21:20
74e7900

Proofs you can hand out, providers you can switch between.

v0.7.0 deepens the audit trail and broadens where bide runs. A proof now reveals only what it proves, every provider adapter speaks its API exactly, and the journal returns byte for byte what it recorded, on every store and across every process.

  • Private, precise proofs. Every journal record and event carries its own salt, so a proof bundle discloses the record it proves and nothing about its neighbours. Signed tree heads commit to their tree's kind and run, evidence packages are sealed and checked field by field, and delegated grants attenuate strictly (issuer, expiry, and every scope key).
  • Earned authority with a current-grant ledger. VerifyCurrentGrant confirms a grant is the latest one issued, against a ledger head the verifier has seen, so a demotion takes effect everywhere.
  • Exact replay. A resumed run sends the model the same bytes the live run did, on MemStore, SQLite, and Postgres alike. agent/durabletest checks any store against this contract.
  • Provider adapters, refined. Requests match each provider's API: grouped Gemini tool results and schemas, Gemini thought signatures and Anthropic redacted thinking carried round-trip, unique tool-call ids on every turn, and max_completion_tokens for OpenAI reasoning models. Quota exhaustion is reported as ErrQuotaExhausted rather than retried, and a turn ends only on the provider's own end-of-turn signal.
  • Durable governance at scale. Governed event logs append idempotently across Redis, SQLite, and Postgres, migrate safely under concurrent opens, and work on Redis Cluster. Each quorum has its own name, and a tie is never agreement.
  • Sagas and sub-agents. Rollback reaches every call a saga started, including ones cut off mid-flight and those inside sub-agents. A pause inside a sub-agent names the root run to continue, a pause in one parallel tool lets its siblings finish, and Recover knows a finished rollback is done.
  • Sessions and flows. Sessions stay consistent across several workers, with each turn tied to its message. plan rejects graph shapes it cannot run as declared, and a run always resumes under the flow it started with.
  • Privacy by default. Traces and ToolLog record an error's category unless content capture is on, URL credentials are redacted from tool errors before they are journaled (WithToolErrorRedactor scrubs more), and signers and adapters never print their keys.
  • Verification you can re-run. Fuzz targets for every parser and verifier run in CI alongside govulncheck, a Benchmark workflow reproduces the published numbers on a standard runner, and How bide is verified describes the discipline behind each guarantee.

API changes:

  • Audit formats move to bide.audit.sth.v4 with versioned, salted leaves; re-anchor heads made with earlier versions. Record gains Salt, and EventLog.Prove returns EventInclusion.
  • VerifyRun, VerifyDelegationChain, SignAbsenceRoot, the absence proofs, EvidencePackage (Seal, WithConsistencyFrom), and earned authority (ProveCurrentGrant, VerifyCurrentGrant) have new signatures.
  • govern.EventLog.Append takes an append id, govern.Quorum takes a name, and bide-audit verify-quorum takes -name.
  • agent.MemWaker.Start returns a channel that closes when it stops, and a model turn with a missing or reused tool-call id is an error.
  • Postgres keeps its journal in the bide_steps table.

Install:

go get github.com/bide-ai/bide@v0.7.0

Docs: https://bide-ai.com

bide v0.6.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 06:07
4b7dcbe

At-most-once everywhere you author work.

v0.6.0 carries bide's side-effect guarantee to every place work gets written: plain-Go steps, fan-out tasks, tool middleware, sessions, and MCP tools, with cleaner cancellation and streaming along the way.

  • Durable steps run at most once. agent.Step and Parallel tasks get the same attempt-marker protection as tool calls and plan nodes. Declare a step safe to re-run with StepSafety (or Task.Safety).
  • Safety-aware tool middleware. ToolRetry retries only retry-safe tools and ToolCache caches only read-only ones. The agent also guarantees that a tool that isn't retry-safe runs once per call, whatever the middleware does. Middleware reads a call's safety with agent.ToolSafety.
  • Idempotent conversations. Session.SendOnce answers each inbound message once per key, so redelivered webhooks never open a second turn. Sessions stay consistent across several workers.
  • Per-run token budgets. Agent.WithTokenBudget is rebuilt from the journal, so a budget holds across a crash and resume. Each model call's usage is recorded with its turn.
  • Consistent usage accounting. agent.Usage counts each token once on every provider. TotalInputTokens and TotalTokens give the sums, and cost is comparable across Anthropic, OpenAI, and Gemini.
  • Clean cancellation. A cancelled tool call is treated as an unknown outcome rather than a failure. A cancelled run stops at the next turn, and a parent waits for its sub-agents.
  • Streaming. agent.NewStreamFunc and Stream.Close release a model response as soon as the caller stops reading. A response that ends before the turn finishes is reported as ErrIncompleteResponse. model/modeltest checks any adapter against the streaming contract.
  • MCP. Tool annotations apply for servers you trust, with mcp.TrustAnnotations(). Tools from other servers are treated as side effects.
  • Evals. Every eval samples the model afresh, even over a store kept between evals.

API changes:

  • middleware.TokenBudget is replaced by Agent.WithTokenBudget.
  • agent.Step treats a step as a side effect unless StepSafety says otherwise, and Parallel requires unique task names.
  • Usage.InputTokens is the uncached input on every provider.
  • eval.Run returns (Report, error).
  • mcp.Tools applies annotations only with TrustAnnotations().

Install:

go get github.com/bide-ai/bide@v0.6.0

Docs: https://bide-ai.com

bide v0.5.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 04:15

Tighter guarantees for multi-process deployments.

v0.5.0 extends bide's durability and audit guarantees to more of the ways a real deployment runs: several nodes sharing one journal or one governed event log, handoffs between them, and streamed runs.

  • At-most-once across overlapping drivers. When two nodes briefly drive the same run, as can happen around a lease handoff, a non-idempotent step is claimed exclusively before it runs, so it still fires at most once. Leasing remains the efficiency mechanism; the claim is the safety one.
  • Governed state shared live across processes. Governors sharing an event log now fold in each other's events as they act, not only on restart, and every governed action records its position in the log, so an auditor can check each recorded state digest by replaying the log through that position.
  • Strict per-entity order on Postgres. The Postgres event log keeps each entity's events in a fixed order under concurrent writers.
  • Hedged streaming. With backup models configured, a streaming caller receives exactly the response the run records.
  • Faithful in-memory store. MemStore returns independent copies in the same form the SQLite and Postgres stores do.

API changes in govern:

  • EventLog.Append returns the event's position, and Events takes a starting position. govern/eventlogtest checks any implementation against the contract.
  • Applier.Apply returns Applied{State, Position}, and FederatedApplier.Apply returns FedApplied. Log-backed governors gain Sync(ctx).
  • The Postgres event log uses a new governed_events table. Redis event-log streams should be written only through the adapter.

Install:

go get github.com/bide-ai/bide@v0.5.0

Docs: https://bide-ai.com

bide v0.4.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 02:42

m-of-n human approval, provable offline.

A tool can now require k signed approvals from a named set of n approvers before it runs. The run pauses durably until enough decisions arrive, proceeds at k, and denies once k can no longer be reached.

refund := agent.Func("refund", "refund the order",
    agent.Safety{Approval: &agent.ApprovalPolicy{Need: 2, Approvers: []string{"ops", "finance", "risk"}}},
    doRefund)
a := agent.New(model, store, refund).WithApproverVerifiers(keysByApprover)

// each approver signs the exact call they were shown
sig, _ := signer.Sign(agent.ApprovalDecisionBytes(pend.Subject(), "finance", true))
agent.ApproveAs(ctx, store, pend.RunID, pend.ToolUseID, "finance", true, sig)

Each signature covers the call itself (tool and arguments), so it approves that call and nothing else. A forged or mistaken decision is ignored and does not lock its approver out. Afterwards, audit.VerifyApprovals, or bide-audit verify-approvals for auditors who do not write Go, checks from a portable evidence file that k named approvers signed off on this call before it ran, under the expected policy, and that the evidence leaves no decision out.

A runnable walk-through, with every actor in its own process: examples/approval. Guide: Approval.

Fix: re-invoking a finished run could repeat a side effect. Calling a run again after it had finished (a retry after a lost response, a redelivered job, a sub-agent or session turn re-entered on resume) asked the model for another turn, and a repeated tool request under a new call id could run again. A finished run now returns its recorded answer without calling the model. This affects v0.1.0 through v0.3.0; upgrading is recommended. As a result, calling Run on a finished run id with new input returns the original answer; use Session to continue a conversation.

Everything else from v0.3.0 stands.

Install:

go get github.com/bide-ai/bide@v0.4.0

Docs: https://bide-ai.com