Skip to content

bide v0.8.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 07:06
· 86 commits to main since this release
629d47e

Runs that recover themselves, budgets that cover the whole tree.

v0.8.0 makes durable runs self-sustaining and fully accounted for. Recovery runs continuously, every run journals where it started, token budgets and spend reach every sub-agent and every request sent, and each library module now installs with go get.

  • Continuous recovery. agent.RecoverLoop keeps taking over the runs of a crashed holder with no further calls, with its own interval, concurrency and error handler. Leases are held under a per-call token, failed renewals retry well before the lease expires, and a drive that loses its lease stops with ErrLeaseLost. A multi-process test harness on Postgres kills, stalls and restarts workers to check it, and the Postgres stores run every write at read committed, whatever the deployment's default isolation.
  • Runs that know where they started. Each run journals its input and entry point (run:start), readable with agent.RecordedStart, so a Recover callback resumes a run with exactly what it began with. Resume decisions come from the journal: a recorded approval denial stays final, saga rollback compensates by the safety each call recorded, and a plan node re-runs only when it was retry-safe when attempted.
  • Sessions in their own namespace. A session's journal and turns live under "<id>>@session" and "<id>>@turn/<n>", so no run started with Run can share one, session ids may contain /, and agent.IsSessionRun identifies them. Timers inside sub-agents wake under names qualified by the sub-run, so sibling sub-agents can wait on timers of the same name.
  • One budget for the agent tree. WithTokenBudget on a parent covers its sub-agents, and Result.Usage and the new Result.Spend report the whole run, including retried, hedged and failed requests. middleware.CostMeter reports what was spent, a TurnRestarted event marks a retried streamed turn, and WithModelCallHook and WithModel add hooks and per-call routing around every request.
  • Precise model boundaries. Finish reasons are one typed, provider-neutral set (stop, tool_use, length, filtered) with the provider's own value in Finish.Raw, and a cut-off or filtered turn is reported as ErrOutputTruncated or ErrOutputFiltered. Responses are capped (WithMaxResponseBytes), adapters describe themselves through agent.ModelInfoOf, and the provider HTTP kit has its own package, model/provider, for writing adapters.
  • Typed tools, exactly as declared. Tool arguments decode strictly against their schema, schema.For describes precisely what encoding/json decodes, and a call whose outcome is lost mid-flight on a tool that is not retry-safe is recorded as unknown (ErrToolOutcomeUnknown). A call cancelled before its effect ran is re-attempted.
  • MCP with limits. mcp.WithSafety declares retry safety per tool, and WithCallTimeout, WithMaxResultBytes and WithMaxDescriptionBytes bound each server, with 1 MiB and 8 KiB defaults. Tool lists are checked for malformed and duplicate names, and modeltest.ToolNames checks provider name rules.
  • Audit artifacts for tooling. Proofs and evidence use versioned formats with snake_case JSON, signed tree heads carry checked timestamps (audit.CheckTimestamp, WithClockSkew), and bide-audit adds -json and -version with one exit-status scheme: 0 verified, 1 not verified, 2 usage, 3 no verdict, 4 input unusable.
  • Flows as versioned config. plan configs declare "version": 1 with snake_case keys, blocks can be named (Arm.Named, plan.BlockName), steps and joins take a context.Context, and a config safety can lower a node's retry safety, never raise it.
  • Evaluation you can gate on. A metric that cannot score a run leaves it unscored, eval.Compare returns a Comparison with a Gate and tolerances, and eval reports errors instead of panicking.
  • Retrieval across stores. RetrievalName and RetrievalDescription let one agent search several stores, and retrieved context is journaled once per run and kept in every model call, including after resume.
  • Journaling throughput at the v0.7.0 level. Recording a step decodes each record once and tags message parts without a re-encode, with the journal bytes unchanged. On a standard 4-vCPU GitHub runner (AMD EPYC 7763, median of 21), 20,000 runs with 5,000 in flight and 50 ms model calls finish in about 1.05 s (~19,000 runs/s), and the overhead scenario runs ~23,700 runs/s (cmd/bench).
  • Every module installable. govern is its own module, and the stores, mcp, trace, codec/gcf, govern and the governed-event logs are tagged with each release, so they install with go get at the same version as the core. CI checks the doc comment on every exported identifier and compiles every Go block in the README and docs against the current code.

API changes:

  • Journal keys encode the tool-use id (tool:<id>, sub-runs <parent>><id>); runs journaled by v0.7.0 do not resume.
  • ResolveHalt resolves tool calls only; use ResolveStepHalt for a Step.
  • Step names with a reserved engine prefix, run and session ids containing >, and plan step names containing : are ErrConfig, and Recover skips sub-runs.
  • Finish.Reason uses the neutral values stop, tool_use, length and filtered; a truncated or filtered turn fails with ErrOutputTruncated or ErrOutputFiltered, any other reason with ErrStreamProtocol.
  • A reply over the response cap fails with ErrResponseTooLarge, and a tool_use turn with no call fails.
  • Tool arguments for Func, SubAgent and final_answer decode strictly: missing required, unknown, case-variant, duplicate or trailing data is ErrToolArgs.
  • schema.For describes what encoding/json decodes and returns ErrUnsupportedType for kinds JSON cannot carry.
  • RunTyped[T] needs an object T, and RunTypedNative on Anthropic returns ErrConfig.
  • WithRetrieval sends context as a user message before the latest user turn, journaled once per run, and RetrievalTool and WithRetrieval require k >= 1.
  • WithTokenBudget on a parent covers its sub-agents, and Result.Usage and Result.Spend report the whole run.
  • agent.EmitMessage takes the turn's Usage.
  • plan step and join functions (Builder.Step, Join2, Join3, RegisterStep, RegisterJoin2, RegisterJoin3) take a context.Context.
  • Audit proof JSON is snake_case with a format field (bide.audit.proof.v2, bide.audit.evidence.v4), and other formats fail with audit.ErrFormat.
  • plan digests are bide.plan.topology.v2 (flows started under v1 do not resume), config JSON decodes strictly, and a config safety can lower retry safety but not raise it or clear approval.
  • plan configs need "version": 1 and snake_case keys, Topology JSON keys are snake_case, TopologyNode.Kind is TopologyNodeKind, and parse errors wrap ErrConfig.
  • Signed tree heads need a positive timestamp within the clock skew, empty evidence packages fail, and bide-audit input is capped at 256 MiB.
  • ApprovalPolicy.Validate refuses approver ids equal under Unicode normalization and case folding, and invalid UTF-8.
  • middleware.Retry and ToolRetry with n < 0 and Quorum with k above the voter count are ErrConfig, and EventLog.Events returns ErrProtocol on a gap.
  • eval.Judge passes only on an exact PASS, and a failed run never passes.
  • eval.RequiredRuns returns (int, error), and eval.Run returns an error for duplicate metric names.
  • eval.Matches takes a *regexp.Regexp, AgentRunner returns (RunFunc, error), Compare returns (Comparison, error), and GovernanceHeld predicates take a context.
  • The provider HTTP kit moves from agent to model/provider; WithToolResultCodec takes a provider.ToolResultCodec.
  • Finish.Reason is the typed agent.FinishReason.
  • eval Metric.Fn returns (bool, error), MetricStat reports Scored and Unscored, and reports are bide.eval.report.v2.
  • bide-audit exit statuses: 0 verified, 1 not verified, 2 usage, 3 no verdict, 4 input unreadable or unusable.
  • A session's journal and turn runs are "<id>>@session", "<id>>@turn/<n>" and "<id>>@event/<encoded key>"; sessions journaled by v0.7.0 open empty, and Recover skips them.
  • Every run journals run:start; resuming with a different input, a saga through Run (or a run through RunSaga), or SendOnce with a different input on an open turn is ErrConfig.
  • plan attempt markers record retry safety; a node re-runs on resume only if it was retry-safe then and now, and a marker written before v0.8.0 halts.
  • A recorded approval denial is final even if the tool's gate is later removed, loosened or made m-of-n.
  • Saga rollback treats a completed call as a write unless its result records it ran ReadOnly, and reports calls to unregistered tools as uncompensated.
  • plan.Retryable and the config safety "retryable" are removed; use plan.Idempotent and "idempotent".
  • govern is its own module: add github.com/bide-ai/bide/govern with go get.

Install:

go get github.com/bide-ai/bide@v0.8.0

Library modules, at the same version:

go get github.com/bide-ai/bide/store/sqlite@v0.8.0
go get github.com/bide-ai/bide/store/postgres@v0.8.0
go get github.com/bide-ai/bide/mcp@v0.8.0
go get github.com/bide-ai/bide/trace@v0.8.0
go get github.com/bide-ai/bide/codec/gcf@v0.8.0
go get github.com/bide-ai/bide/govern@v0.8.0
go get github.com/bide-ai/bide/govern/sqlitelog@v0.8.0
go get github.com/bide-ai/bide/govern/redislog@v0.8.0
go get github.com/bide-ai/bide/govern/postgreslog@v0.8.0

Docs: https://bide-ai.com. The approval guide is now Human approval (human-in-the-loop).