Releases: Egoushka/chargehand
Release list
0.6.1
Two fixes found by the live checks of 0.6. A worker's citation of a service tool's reply now resolves, and a memory
mapping can carry where a run's citations were checked into the retain call's metadata. The breaking change and the
migration notes of 0.6.0 are unchanged. The one schema edit is the description of the locator field in result/v1, and
Chargehand.Contracts 1.3.1-alpha is published with it, because the package embeds the schema text.
Added
- Retain placeholders
{repository},{commit}(12 hex characters) and{locators}(joined with;), so a memory mapping can
send where a run's citations were checked as metadata (profiles/example.jsonand the guide's Hindsight entry do). Hindsight
rewrites a retained item into a sentence without the commit; the metadata and the stored document keep it. A recall
results.texttemplate can read nested fields ({metadata.commit}). Mappings without the new placeholders send what they sent.
Fixed
- A worker's citation of a service tool's reply now resolves. Workers cannot see message ids, so they cited the reply text as a
session_messagelocator and the claim became an open question. Asession_messagelocator that quotes at least 12 characters of a tool's reply in the session now resolves, and the task text says so when the run has services. Nothing changes without services;result/v1changes only in the locator's description.
0.6.0
Breaking: the single-object form of memory in the profile no longer loads. Memory is a list of MCP providers now, and a
profile that still has the object fails with memory is a list now and a pointer to the guide. Move the memory
service's MCP endpoint into mcp_servers and list the provider under memory; Removed below has the migration, and a
profile without memory needs nothing.
Goal 0.6 (roadmap): runs use your MCP services and memory. Memory is an ordered list of MCP providers, each
with a declarative mapping from recall, retain and invalidate onto the server's tools; recall asks all of them and
stacks the facts with the name of the memory each came from. With retain on, a run stores only the claims whose
citations resolved at its pinned commit. A preset can give its workers read-only tools from an MCP service, on Claude
Code and on OpenCode, and chargehand extensions check tests the mappings and grants against the servers' real tool
lists. With no server listed a run behaves as before. Memory and services walks
through the setup. The release also closes a hole in the OpenCode runtime: workers could start a checkout's own MCP
servers and call any MCP tool. Project config is now disabled on the servers chargehand starts, and MCP tools are denied
unless a preset grants them (Security; restart an OpenCode server you started earlier).
There is no 0.5.0: goal 0.5's code shipped in 0.4.0, and what remains of its done-when is the maintainer's own use of
/chargehand:change on real tasks, which is still open. The one schema change is additive: preset/v1 gains the
optional services field, and Chargehand.Contracts 1.3.0-alpha is published with it.
Security
- OpenCode workers can no longer reach MCP servers. A preset's leading
* * allowreached the tools of any server
registered at the location, including write-shaped ones and OpenCode's MCP resource tools, and a checkout's own
opencode.jsoncould register a server whose command OpenCode started when a session was created there. Every
OpenCode session's rules now end with*_* * deny. The server chargehand starts andscripts/opencode-serve.shset
OPENCODE_DISABLE_PROJECT_CONFIG=1andOPENCODE_CONFIG_PROJECT_DISABLE=1, so a checkout'sopencode.json,
.opencode,AGENTS.mdand project skills are no longer read (workers stop receiving the checkout'sAGENTS.md).
Restart a server you started earlier with the script, and set both variables on one you start another way. The Claude
Code runtime was not affected (ADR 0034).
Added
mcp_serversin the profile (ADR 0034): MCP servers by name, over Streamable HTTP, the older SSE transport
("transport": "sse";auto, the default, tries Streamable HTTP and then SSE, and means Streamable HTTP for a worker's
runtime) or stdio, with{secret:item}in header and environment values. Memory and a preset'sservicesboth read
them, over one connection per server.- Memory from any MCP memory server, several at once (ADR 0034).
memoryis a list of providers; each names an
mcp_serversentry and maps recall, retain and invalidate onto that server's tools: argument templates with{query},
{namespace},{max_facts},{text}and the like, and aresultsmapping that says where the facts are in the answer
(path,id, andtextas a field, a template such as{date}: {summary}, or an ordered list of these). Recall reads
a tool's structured content when the path finds an array there and its text block otherwise. An entry without a retain
tool is recall-only, as an archive such as Chronicle is, andnamespacedefaults to the entry's name. A mapping that
names an unknown server or placeholder, or misses its recall tool, fails when the profile loads. serviceson a preset's node kind (preset/v1, an additive field; ADR 0034): a server frommcp_serversand the tool
names workers may call, exact or with*globs, never a whole server. At the start of a run chargehand connects,
lists the server's tools and grants the ones named. A server, secret or tool that does not resolve is dropped, not fatal:
chargehand showprints oneserviceline each, with the span tagschargehand.service.<name>.grantedand.issues.
as_sent.tools_sha256covers the granted tools and is unchanged when there are none. No shipped preset lists services.- Claude Code workers get the granted services (ADR 0020 addendum): the servers go in a private
--mcp-configfile, mode 0600
in a fresh 0700 directory under the system temp directory, removed when the turn's process exits (also on failure or
interrupt) and never on the command line or in a log;--allowedToolsnames exactly the granted tools, and the server's
other tools are disallowed so they do not cost tokens. A granted server the CLI reports as not connected is an issue on
the run (not_connected: failedinchargehand show), not a silent gap. - OpenCode workers get the granted services (ADR 0034): before a run's first session at a location chargehand registers each
granted server there (PUT /api/experimental/mcp/{name}), waits untilGET /api/mcpsaysconnected, and ends the
session's rules with*_* * denyand one<server>_<tool> * allowper granted tool, so only those tools are callable
and a server's other tools stay refused. The registration is named for the server and a per-process keyed hash of its
config, so runs with other credentials do not change each other's server; it is shared by every run at the location
that uses the same config, held by count, and removed when the last run ends (completed, failed or cancelled). A
server that does not connect (failed,needs_auth, no answer in 30 s) is dropped, removed and reported on the run as
not_connected: <status>; the run goes on without it. Presets that deny*still cannot reach a service on OpenCode
(workers use itsexecutetool); no shipped preset lists services. chargehand extensions check [--preset <name>] [--probe <query>](ADR 0034): connects everymcp_serversentry and lists
its tools, then checks each memory mapping (the tool exists, every argument name is a property of its input schema,
every required argument is set) and each preset'sservices(every named tool or glob matches a listed tool). It prints
one line per item,okor the problem with; action:and what to do, and reads every preset inpresets/unless
--presetnames one.--probealso runs one real recall per memory and prints how many facts came back, never the
facts. Exit 0 when nothing is wrong, 1 on a problem, 2 for a usage error or a profile that does not load. A line never
holds a URL, a credential or an argument value.- The guide page Memory and services: the
mcp_serverstransports, the Hindsight-through-a-gateway
and Chronicle entries side by side, retain, preset services and what each runtime does with them, the check command,
the migration from the oldmemoryobject, and the live checks made on 2026-09-29.
Changed
- A
memorylist in the profile is now read (ADR 0034); until now it loaded and did nothing.run,serveandmcpbuild
one memory stack from it: each entry becomes a source that calls its server's mapped tools over the same MCP
connections a preset'sservicesuse, in list order, with the entry's limits,retainandretain_tags. A profile
copied fromprofiles/example.jsonnow tries its memory servers when a run recalls; a server or secret that does not
resolve skips that memory, andchargehand showsays why. - Recall asks every memory at once and labels each fact with the memory it came from, and the prompt header says so:
- [hindsight] Deploys go through GitOps.The chain block for recalled text is namedmemory/recall/<name>, one per
memory that contributed, instead ofmemory/recall. Each memory keeps at most 10 facts and 4000 characters, cuts a fact
at 600 characters, shows each fact as one line and gets 10 seconds by default (timeout_seconds); a fact two memories
return appears once with both names. One recall per run and retain off by default are unchanged. A memory that throws,
times out or answers what its mapping cannot read is skipped on its own; the others still contribute. chargehand showprints one line per memory with what it recalled and retained, or why it was skipped, from a new
optionalextensionsreport in the run record. The span tags for memory are per source
(chargehand.memory.<name>.recalledand.error) instead ofchargehand.memory.recalledand
chargehand.memory.error.result/v1is unchanged.- Retain keeps less, and what it keeps says where it was checked. With
retainon, a run used to store the request text,
the summary and every claim. It now stores only the claims that cite at least onefileorcommitentry that resolved
at the run's pinned commit, each with its locators, under a first line that names the repository and the commit:
Repository: github.com/example/proj, commit 0123456789ab (citations checked at this commit), then
- <claim> [src/Api/Startup.cs:41-58] (confidence 0.90). No longer retained: the request text, the summary, claims
that rest only on caller inputs, URLs or the session, claims whose citations did not resolve, claims whose text the
error-text scrubber would change, and everything from a run without a repository (a draft run has no commit, so it
retains nothing). The repository is theoriginURL without scheme, user information, port and.git, or the
directory name when there is noorigin. A run that retained nothing says why inchargehand show
(no commit,no claim qualified). There is no confidence floor. - A command secret source that runs longer than 15 s is killed and the next source tried; the error says when one timed out.
- T...
0.4.1
Fixes found in use of 0.4.0 and the first pieces of goal 0.5, and the first version the MCP Registry can list, under
the name io.github.Egoushka/chargehand. Intake now sees the text of caller inputs, within a cap per input and a cap
in total; a failed worker session and a Claude Code error result always carry a cause; memory that times out or
answers garbage is skipped instead of failing the run; and the private-terms hook works in a linked worktree. No schema
changed, so Chargehand.Contracts stays at 1.2.0-alpha.
Added
- The release workflow lists each new version in the MCP Registry (ADR 0027). A
mcp-registryjob runs afterrelease
whenNUGET_USERis set: it stamps a copy of.mcp/server.json, waits until nuget.org serves the package README with
the ownership line, and publishes with the registry's GitHub OIDC login. It holds no write token and needs no new secret.
Changed
- The MCP Registry name is
io.github.Egoushka/chargehand, with the owner spelled as GitHub spells it: the registry
matches the namespace and the README'smcp-nameline case-sensitively. TheChargehand0.4.0 package has the
lower-case line and cannot change, so 0.4.1 is the first version the registry can list.
Fixed
- Intake saw a caller input's id and kind but not its text, so a run whose request carried its goal, diff and test
output asinputscould stop withask, asking for the inputs it had been sent. Intake now reads each
input's id, kind, size and the first 2000 characters of its text, and the prompt says when an input was cut; the
worker still gets every input whole. - A worker session that ended
failedwith no assistant message and no error text (OpenCode drops the session before
any model call when the model is one its server does not declare) returned a bare "worker ended failed" with a
generic action, and the reason was only in the server's log. The error now says the worker runtime ended the session
before any model call, and itsactionpoints at the runtime's log (for OpenCode, the server's) and names the model
the run asked for, to check against the runtime's models and the profile'smodelsmap (or says no model was mapped). - A Claude Code turn that ended in an error result with no
resulttext (error_max_turns,error_during_executionand
error_max_budget_usdcarryerrorsinstead) failed with no reason and was reported in OpenCode's terms, as a session
ended before any model call. The error now names the result's subtype, turn count anderrors. - A request with many
inputsstill grew intake's prompt without bound, up to the request size limit. Intake now reads
at most 8000 characters of input text in all, in request order; an input after that is listed with its id, kind and
size and a note that its text is left out here, and the worker still gets every input whole. - The
.private-termshooks work in a linked worktree. The list is gitignored, so a worktree never had a copy and every
commit there aborted with ".private-termsis missing"; the check now falls back to the main worktree's list. A
worktree's own file still wins, and with no list in either place the commit is still blocked. - Memory failed open only on
HttpRequestException. A provider timing out (TaskCanceledException, whichHttpClient
raises when its own timeout elapses) ended the run with an exception and no result, and aJsonExceptionor
IOExceptionfrom a provider failed it. Recall and retain now skip the provider on any exception except the caller's
own cancellation, and the run goes on without it.
0.4.0
Goal 0.4 (roadmap): chargehand runs with nothing configured. No profile file is needed, the worker
runtime is named or found on the machine, credentials come from environment variables or the CLI's own login, and the
orchestrator ships as a dnx tool that speaks MCP over stdio. The release now publishes that tool, Chargehand, to
nuget.org next to Chargehand.Contracts. The first pieces of goal 0.5 (the review preset, the Claude Code plugin and
the change skill) are on main and in this tag; the plugin's server needs the published package.
Added
chargehand mcpserves theorchestratetool over stdio, so an MCP client starts it itself; its child processes get
a closed stdin, and a caller receives the run id before its own timeout can fire.- The orchestrator packs as a
dnxtool,Chargehand, with an MCP server manifest (.mcp/server.json, ADR 0027),
packed and started in CI by a smoke test. Prompts and presets are read from the install, and the run log goes to a
per-user default when no profile names one. - Runtime detection without a profile block: Claude Code connects on its own (and on the CLI's own login when no
credential variable is set), and an own OpenCode server is started when the profile has noopencodeblock. OpenCode
is pinned at 2.0.18. - CI fails on a breaking change to a published schema since the latest release tag (additive changes only).
- The read-only
reviewpreset, thechargehandClaude Code plugin and its marketplace entry, and the change skill
(/chargehand:change, one prompt to a reviewed change), with an end-to-end check and a user guide underdocs/guide.
A workflow that has chargehand review its own pull requests exists and stays off untilSELF_REVIEW=true. - ADR 0033, the CI policy: one required
ci-gatecheck, prose-only pull requests skip the build jobs, CodeQL and
Scorecard run weekly and on demand instead of on every pull request or push, and auto-merge is on.
Changed
- Profile is fully optional (ADR 0026):
Profile.Loaddefaults an absent file rather than throwing;worker_root,
default_preset,intake_modelandpriceseach have a default (a fixed directory outside$HOME,cheap,
unset, and empty).secret_storeis replaced by an orderedsecretslist (env, then command templates as argv
arrays, first success wins) — a profile still carrying"secret_store": "keychain"needs a one-line migration
to"secrets": [{"env": true}, {"command": ["security", "find-generic-password", "-s", "{item}", "-w"]}](see
profiles/example.json). - The worker runtime is chosen by a
RuntimeSelector: a profile'sruntimefield orCHARGEHAND_RUNTIMEwins,
then the profile's only runtime block (opencodeorclaude_code), so an existing profile keeps its runtime
(ADR 0032); otherwisePATHis probed for a known agent CLI. No agent CLI found isruntime_unavailable; more than one found
is a newruntime_ambiguouserror naming every CLI seen — no silent priority order. IPriceTable.PriceUsdreturnsdecimal?: an unpriced model's cost is unknown, not a silent$0. The run's USD
cap cannot fire on an unpriced model; the per-node-kind token budgets remain the real guardrail.result/v1's
usage.usdmay now benullfor the same reason.ClaudeCodeWorkerRuntimeandOpenCodeWorkerRuntimeaccept an unset model (NodeSpec.Model,GenerateAsync) and
fall back to the runtime's own default instead of requiring one.
Fixed
- Every
result/v1error now names a concreteaction, andrepository_not_allowedsays how to fix it; a run with no
repository_rootsallows the directory it was launched in. - With no profile, or a
modelsmap that does not name a preset's placeholder model (provider/worker-model,
provider/small-model), the placeholder reached the agent CLI and every run failed. An unmapped placeholder is now
unset, so the runtime uses its own default model, labelledautoin the prompt chain as for intake.
Realprovider/modelids in a preset still pass through. - A call the runtime reports no model for (Claude Code's compactions; every call on the runtime's default model) cost
a silent$0. It is now priced at the node's model, and with no model on either itsusage.usdisnull. - A worker that ended
failedreturned "worker ended failed" and nothing else: the provider's reason was in the
session messages and never reached the result. The message now carries it (worker ended failed: There's an issue with the selected model …), scrubbed and cut at 300 characters, and every node failure — failed, rate limited,
over budget, past its deadline, no valid result — carries anactionsaying what to change or where to look
(chargehand show <run id>). - A worker that wrote
status: failedin its own result block (a change request on a read-only preset, which runs as
an answer when intake's action is not in the preset) returned a failed result with noerror, against ADR 0022. It
now carriesinternalwith "the worker reported failed:" and an action. - The release workflow published only
Chargehand.Contractsto nuget.org, so theChargehanddnx tool package the README,
the plugin and the MCP Registry manifest name was never available. It now packs and pushes the tool too, each package
only when its version is new there. - Prompt CI blocked every pull request that adds a preset with its eval cell as uncovered: the gate reads cells from
main, which does not have the new cell yet. A file the base lacks now passes when a cell in the change's own
evals/cells.jsonnames it (--change-cells-file); tolerances and existing files stay on main's cells. With
--allow-uncovered, the status said "no prompt or preset change"; it now names the files that passed without a cell. - Prompt CI crashed with an unhandled 404 when an eval cell's Langfuse dataset did not exist yet (a new cell, before
its firsteval push). A missing dataset now has no items, so the gate blocks with "0 items; the gate needs at
least 8" instead.
Security
- The
.private-termsdenylist fails closed. A malformed pattern madegit grepexit 128, which thepre-commit
hook read as "no match" and let the commit through; the commit is now blocked with the error. Thecommit-msghook
checks the message against the same list, except your ownSigned-off-byand the diffcommit -vadds
(scripts/check-private-terms.sh). - Prompt CI evaluates the commit its run was approved for (the event's head, passed as
HEAD) instead of reading the
pull request's head when the gate job starts, so a push between an approval and the job no longer runs
unreviewed prompts on the eval runner. Scrubno longer redacts ordinary words that merely contain a key-like substring (task-spec,disk-cache,
a stray "user:") — thesk-pattern needed a left boundary. It now also catches shapes 0.2.2 missed: LiteLLM's
"Key Hash (Token) =", compound identifiers (team_member,user_id,organization,user_api_key_alias),
camelCaseapiKey,Authorization: Basic, and spend figures in scientific notation. Fixes an ordering bug where
atoken: Bearer <secret>message redacted the word "Bearer" instead of the secret.
0.3.0
Goal 0.3 (roadmap): every change is tracked, released and checked. Its substance shipped across
0.2.2's patches (the release pipeline, quality gates, repository hygiene); this bump closes the goal and moves the
roadmap to 0.4.
0.2.2
Security
- A failed result no longer repeats the model gateway's error text verbatim: keys, key aliases, bearer tokens and
spend figures are redacted before a message reachesresult/v1or the run log. A refusal over a spend budget says
what to do inerror.action.
Added
ROADMAP.md;scripts/check.sh, the one check before a push; a Claude Code hook that asks before an edit to a
published schema major.- OpenSSF Scorecard, dependency review on pull requests, CodeQL, and a SonarQube Cloud job that runs once the project
is connected; badges in the README.
Changed
- A tag
v<Version>also creates the GitHub Release, with the version's section of this changelog as its notes, and
publishesChargehand.Contractsto nuget.org when its version is new there. Publishing waits for an approval in the
releaseenvironment; a tag whose version has no section here fails before anything is published, and CI fails a
version bump that comes without one. - The Prompt CI runner ADR is now ADR 0025 (two ADRs had number 0022); ADRs 0024 and 0025 are accepted.
TRADEMARK.mdsays how the name may be used, and contributions are signed off (DCO). - The build runs the .NET analyzers at
10.0-recommended. Parsing and formatting no longer depend on the machine's
locale: a hand score such as0.8parses where the decimal separator is a comma.
0.2.1
0.2.0
The server runs on a private network
chargehand serve can bind a private-network address: http.listen sets it and http.allowed_hosts names the host
clients use; without allowed hosts it refuses to start beyond loopback. A Dockerfile builds the server with the
Claude Code runtime, and a tag v<Version> publishes ghcr.io/<owner>/chargehand:<Version>. See ADR 0024.
Prompt CI runs on its own
A pull request that changes prompts/ or presets/ is gated on a self-hosted runner, not by a manual run of
scripts/prompt-ci.sh. A fork's pull request or a preset change waits for an approval in the prompt-ci-review
environment. The runner holds one eval profile per runtime and a default; a prompt-ci:<runtime> label picks another
when someone with write access adds it. The manual run still works. See ADR 0025.
Added
result/v1has an optionalerroron failed results (ADR 0022): a fixedcode, themessage,retryable, and
anactionwhen there is something to do, such as the command that starts the OpenCode server. Clients branch on
the code instead of parsingsummary. The schema stays v1;Chargehand.Contractsis 1.2.0-alpha with the
ResultErrorrecord and theErrorCodeenum.
Changed
- Preset blocks
default0.5.0,cheap,thoroughandstrict0.4.0 say "read the files that answer the task"
instead of "read every file the task needs". Worker prompt 0.3.0 keeps the completeness rules. In the phase 3
benchmark the old wording cost 28% more than a plain session through extra reading; the new one costs the same
and stays more complete (docs/benchmarks.md). - Worker prompt 0.3.0 asks for the whole task, one claim per item and a citation for every file relied on, and
no longer caps the summary at 120 words. Preset blocksdefault0.4.0,cheap,thoroughandstrict0.3.0
say "read every file the task needs" instead of "prefer the smallest set of files". The phase 3 blind verdict
failed on completeness in 3 of 3 pairs at equal exploration: the worker read files it then left out, and merged
several services into one claim. - A worker interrupted at its token budget (
budget.max_input_tokens) gets one turn to answer from what it has
read, with room for three calls at the last call's context on top of what it spent; the result lists the stop in
open_questions. It used to fail with nothing. The USD cap still ends a node without that turn (ADR 0010). - A request's repository may sit anywhere under the profile's new
repository_roots(default:worker_root;/
allows any). The worker reads a clone of it at the pinned commit underworker_root, reused per source and commit,
so the source's uncommitted and ignored files never reach it and the worker stays outside the home. A short commit
hash is enough. A checkout that tracks a file the preset denies reading is still refused (ADR 0023). chargehand runprints a failedresult/v1witherrorwhen it cannot connect to the runtime (server not
running, binary missing, another version), instead of ending with an unhandled exception.- The eval gate retries an arm whose result has code
rate_limited, instead of matching "rate limit" in its summary.
Prompt CI calibration
cheap/worker's items pinned a checkout that tracks encrypted env files the cheap preset denies reading, so since
workers refuse such a checkout (0.1.0) every arm of its A/A failed at $0. The items now pin a commit of that checkout
without those files; every reference file is unchanged. The A/A there, the first since workers lost the shell, 12
items: quality -0.046 (t -0.63), cost +13% (t 1.82), pass, $0.11; 5 items differ, by 0.75, 0.27, 0.25, 0.20 and 0.02.
Claims per item: 6.25 against 5.97 in the runs that seeded reference_claims, 20 of 24 arms within their range, so
reference_claims stays.
0.1.0
First usable release. chargehand turns a request into a Task Spec, runs it on one or more coding-agent sessions, and
returns result/v1 with evidence behind every claim. You call it from the CLI, over HTTP, or as an MCP tool. It
covers roadmap phases 3 to 5 (v0, v1, v2). Benchmark and exit-check numbers live in
docs/benchmarks.md.
Added
Runs and actions
chargehand runreadsrequest/v1, runs intake to get a Task Spec, and returnsresult/v1.- Intake picks one action.
answerruns one worker session.splitruns 2–4 read-only subtasks as a task graph.
denyreturns statusdeniedwith a reason and an unblock condition.askreturnsneeds_inputwith questions.
improvereturnsneeds_inputwith an improved request and an inline diff artifact. When the preset does not allow
the chosen action, the run falls back toanswer; the run log records both. - Split runs (ADR 0017) execute at most 2 nodes at once and pass upstream contracts to dependents; a failed node stops
only its dependents. Later nodes fork the first node's session before its first message, so they read its system
prefix from cache. Node contracts merge deterministically, with evidence ids prefixed by the node id. - The worker node fixes its instruction entries before the first prompt. A watcher rejects permission requests and
interrupts above the run cap. Each node gets a 15-minute deadline and one repair turn each for the result schema
and for evidence. - Per-node budgets: the watcher interrupts a node above
budget.max_input_tokensand compacts it mid-turn above
compaction.trigger_tokens(new optionalpreset/v1field). - The evidence resolver checks claims against git at the pinned commit, session messages, caller inputs, URLs seen,
and the session diff. Claims that do not resolve move toopen_questions. request/v1gains optionalcontext.approvedto run past a preset's approval thresholds.
Presets
default, pluscheap0.1.0,thorough0.1.0 andstrict0.1.0, all read-only.strictasks for approval above
risklowor an estimate above $0.50.draft0.1.0 (node kinddraft) writes a draft for a program caller from the caller's own inputs, with no
repository and no tools, and returns it as an inline artifact. New optionalpreset/v1fieldcheckout: false.
Callable interface (ADR 0018)
chargehand servehosts HTTP and MCP on 127.0.0.1. Every route requires a bearer key from the secret store
(profilehttp), and bodies cap at 1 MB.POST /v1/runstakesrequest/v1. If the run finishes withinPrefer: wait=N(default 10 s, at most 60) you get
200withresult/v1, otherwise202withrun-status/v1.GET /v1/runs/{id}answers202while running,200when finished,410when the owning process died,404
when unknown.GET /v1/runs/{id}/eventsstreamsaccepted,started,intake,node_started,node_finished
andrun_finished.- Runs outlive the request and execute one at a time. The server holds at most 10 unfinished runs, then answers 429.
- MCP at
/v1/mcp(Streamable HTTP, MCP 2026-07-28 with hybrid sessions, C# SDK 2.2.0) exposes toolorchestrate,
withinputSchemarequest/v1andoutputSchemaresult/v1. Long runs become tasks through the tasks extension.
Anaskbecomesinput_required, and the answers resend the request as a child run. - The run log's new
startrecord (request, trace id, owning pid, parent run) makes it the store the CLI and the
server share.showprints runs that are still running or were lost. Concurrent appends take a lock file, and
readers skip a half-written last line. Chargehand.Contracts1.1.0-alpha adds schemarun-status/v1andRunStatus;PromptBlock.Createhashes a caller
block by request/v1's rule.samples/ContentEngineCallbuilds the content engine's call (ADR 0014) fromChargehand.Contractsalone.
Worker runtimes
- OpenCode V2 adapter pinned to 2.0.16: a hand-written client and
IWorkerRuntimeimplementation, with a retry for
theModel unavailablerace while a location boots.scripts/opencode-serve.shstarts the orchestrator's own
server;profiles/opencode.example.jsonconfigures it. - Claude Code adapter (ADR 0020,
Chargehand.ClaudeCode) runsclaude -pwith stream-json, one process per turn,
--bareanddontAsk, and translates preset rules to--tools,--allowedToolsand--disallowedTools. Profile
claude_codeselects it instead ofopencode, which becomes optional. It takesversion,binary, and exactly
one ofapi_key_secret(per-token API billing) oroauth_token_secret(aclaude setup-tokensubscription
token). Optionalbase_urlroutes workers through an Anthropic-compatible gateway.
Prompts, evals and routing
- Prompt registry in
prompts/with SemVer front matter and normalised sha256. Every call records its
prompt_chain;chargehand prompts syncmirrors blocks into Langfuse prompt management. - Prompt CI (ADR 0019):
chargehand eval seed|push|gatewith cellscheap/worker,draft/draftandintakein
evals/cells.json, items in Langfuse datasets, and deterministic scores. It runs base and change in pairs and gates
on quality tolerance T (0.10) and cost tolerance C (+15%, or +30% forcheap/worker), with a declared-trade
override. Results land as Langfuse dataset runs and scores. scripts/prompt-ci.sh <pr>runs Prompt CI on the owner's machine and posts commit statusprompt-ci.
.github/workflows/prompt-ci.ymlmarks pull requests that change no prompt or preset.chargehand routesprints a routing report per preset, node kind and model (runs, score, tokens, cache rate, cost,
latency) and only suggests changes.chargehand scorerecords a hand score for a run.
Memory
IMemoryProvider(recall, retain, invalidate; scope is backend plus namespace) with an adapter for a self-hosted
Hindsight service (HTTP API 0.10.0). With profilememoryset, a run recalls facts once and appends them to each
node's prompt as unverified context (chain blockmemory/recall, sourceruntime). A failed recall leaves the run
without them.retainstores completed runs and stays off by default.
Observability
- JSONL run log with tokens, cache rate, own-table cost and latency per call.
chargehand showsummarises a run;
chargehand reconcilejoins calls to exported gateway spend rows by model, token counts and time. chargehand cache <run>reports cache reads, writes and hit rate per call. Per node, it names the first instruction
entry, prompt block or as-sent field that broke a prefix that should have been shared, plus a checklist for breakers
outside the prompt chain. Call records carry the fork parent and hashed instruction entries.- OTLP traces (run → intake → node → call) carry
chargehand.prompt_chainand the OpenCode session id. telemetry.usage_on_spans(ADR 0021) puts Langfuse usage and cost on call spans that no gateway records, such as
Claude Code on a subscription. Off by default; with a LiteLLM gateway, ADR 0012 still applies.
Changed
- Intake prompt 0.2.0 says when to split. 0.3.0 asks only when a required fact is missing and the caller's inputs do
not supply it; 0.2.0 asked a program caller for facts it had already sent (1 of 2 A/A runs). - Presets
default0.4.0,cheap0.2.0 andthorough0.2.0 drop approval thresholds. Intake's estimate is
uncalibrated (ADR 0005): in the phase 4 benchmark it guessed up to $0.35 for runs that cost about $0.01 and stopped
one withneeds_input.strictkeeps its thresholds. - Preset
default0.2.0 allowsrg(with--predenied). Worker prompt 0.2.0 makes the final message the JSON block
only. - Prompt CI scores a worker as grounding times completeness. Completeness is claims kept over the item's
reference_claims(new optional field in an eval item'sexpected), capped at 1;eval seedproposes the count
from the seeding run. Grounding alone passed #5, which cited the same code in fewer, wider claims. cheap/workerdrops the phase 3 reference question (now 12 items). Atcheap's 400k-token node budget it failed
in about half its runs, and one flip moved an A/A's mean by up to 0.08.- The orchestrator computes an inline artifact's sha256 itself and caps its content at 64 KiB.
Fixed
- A run that throws (bad checkout, unknown preset, no valid Task Spec) ends with a failed
result/v1and a run record
instead of an exception. - Caller inputs appear in the task text as
- id "<id>" (<kind>): <text>. The old[id]form led workers to cite
[id], which matched no input. A failedinputreference now lists the ids that exist, and intake sees each
input's id and kind. - OpenCode's stateless generate (intake) retries one 503, as ADR 0011 prescribes for the masking proxy.
- Prompt CI retries an arm that hits a rate limit (after 15, 30 and 60 s) and stops the gate if the limit persists,
instead of scoring the item 0. - Prompt CI stops the gate (status
error) when a worker or draft arm fails with zero usage, meaning a refusal before
any model call. Such arms used to score 0 on both sides and pass as no change. Intake arms report no usage and stay
exempt.
Security
chargehand serverequires its bearer key on loopback too, since any local process can reach the port. It accepts
only a loopback Host header (against DNS rebinding) and sends no CORS headers.- Prompt CI never builds or runs a pull request's code. The runner uses the trusted checkout's build, and the pull
request contributes onlyprompts/andpresets/. Those still steer a worker, and a preset can grant tools, so a
fork's pull request or any preset change runs only after the owner reads the diff (--reviewed). A symbolic link
among them stops the run. GitHub holds no model, gateway, traci...