graph-agents-cli 0.3.0
0.3.0 lets agents call other agents for the user they serve, over A2A.
Each agent knows the user and the agent in between, a person's approvals stay with that
person, peer add declares the agents one asks, and graph-agents-cli system checks, wires
and deploys several projects as one. It also keeps A2A tasks in Postgres so that replicas
share them, adds structured final answers (an agent that answers in a JSON shape the
project declares), reasoning effort and the Responses API for OpenAI-API models, a
documentation site, skill rules found with the SkillOpt experiment and the benchmark that
measured them, and fixes. What an upgrade from 0.2.0 changes, and the order to do it in, is
in Upgrading projects; the guide to
the new features is Agents calling agents. Parked
medium- and low-priority issues are listed in KNOWN_ISSUES.md.
Breaking changes and migration
Each change below can need an edit in an existing project; the upgrading guide's
0.2 to 0.3 section lists every other
change in behaviour.
jwtreads the RFC 8693actclaim. A token carrying it is an agent's for the user,
refused (403) untilAUTH_ALLOWED_ACTORSlists the agent; setAUTH_JWT_ACTOR_CLAIM=
(empty) to read every token as the user's own, as 0.2 did. Acustompolicy that returns an
invalid id (empty, over 256 characters, or with control characters) now fails the request
with 500 and logs the bug. A custom policy through which other agents forward users'
credentials must setPrincipal.actor(policies/is not upgraded): see the upgrading
guide.- An API whose
auththe project's auth policy can never serve stops the app outside
dev.lintandapi addcheck each API'sauthagainst the project's auth policy:
auth: exchangeorauth: forwardundershared-bearer, andauth: forwardunderjwt
withoutforward_audience, are errors (exit 3): such an API never had a credential to send,
so every call to it failed with "the caller has no credential". OutsideAPP_ENV=devthe
running app now refuses to start with such an API too, as it does forauth: exchange(the
owner's decision of 2026-09-28; alsoauth: forwardunder thelanggraph-serverruntime);
under dev it logs why and starts. Migration: a 0.2 project with such aforwardAPI
stops starting outside dev after the upgrade until the API getsforward_audience, moves to
auth: exchangeor is removed (scaffold upgradenever rewritesapi-policy.yaml, and
lintnames the API).lintandapi addalso note aforwardAPI with
forward_audienceunderjwt("prefer auth: exchange").
Added
- Agents calling agents: the caller's identity. A request another agent presents for a
user is now told apart from the user's own.Principal.idstays the user (the subject);
the newPrincipal.actornames the agent presenting the request (its id, the chain of
agents before it, and its client).jwtreads the RFC 8693actclaim
(AUTH_JWT_ACTOR_CLAIM, nestedactfor earlier agents), the client fromazp/client_id
(AUTH_JWT_CLIENT_CLAIM), and withAUTH_JWT_DIRECT_CLIENTStreats a token with noact
from any other client as that client's; acustompolicy setsactoritself
(actor_from_claimsandkeep_subject_tokenare exported for it). Every policy then goes
through one rule set (finalize_principal): valid ids, at mostAUTH_MAX_DELEGATION_DEPTH
agents (default 3; else 401), only the agentsAUTH_ALLOWED_ACTORSlists (default none;
else 403), and only the rolesAUTH_DELEGATED_ROLESlends. Threads, A2A tasks and approvals
are owned by the subject and the actor: an agent reaches only what it started for that
user, while the user owns everything done for them: with their own token they read,
continue and delete the threads their agents started, decide their approvals, and read,
list and cancel the A2A tasks those agents started for them (the owner's decision of
2026-09-28; continuing or subscribing to such a task stays with its agent). A delegated
principal's roles never read across, administer or decide as arole:approver, and it
never decides an approval (403approval_direct_only: the person decides with their own
credentials). The actor reaches
tools inattributes["@actor"], is recorded with each approval (requester_actor), and is
logged (actor), kept in run records and named in trace metadata.auth dev-tokenmints
such tokens locally (--act, repeatable, and--azp). The database gains
threads.actor,runs.actorandapprovals.requester_actorat startup (ADD COLUMN IF NOT EXISTS); existing rows are direct, and every 0.2 principal is direct, so 0.2 behaviour is
unchanged for them. KI-042 is narrowed (jwtmapsactandazp; one issuer and no
mapping from scopes to permissions remain). - Tools and the model know when another agent asks for the user.
current_caller()
returns the calling agent (Caller.actor,actor_chain,delegated); the new
require_direct_caller()refuses unless the user asks this agent directly;require_owner
still compares the user.require_user_mentionedfollowsA2A_DELEGATED_MENTIONS:origin
(the default) also needs the id in the user's own words the calling agent forwarded, and
refuses when none were forwarded (so, until the A2A client forwards them, a delegated
write the check guards is refused and the user names the record at this agent directly);
refusealways refuses;requestkeeps the 0.2 reading, andlintandapi showpoint it
out when.envor a values file sets it. In a delegated runUntrustedToolResultsfences
each human message the model reads as that agent's (<agent_request from="...">) and adds
one factual note after the system prompt saying an agent wrote the request, with the user's
own words when forwarded;A2A_CALLER_NOTE=offdrops the note. A bad value of either
setting stops startup. Thefaketest model reads the request inside that fence. - Approvals relayed across agents. An approval rule may let named agents deliver the
requester's decision from another agent:decide_with: relayedwithrelayers(actor
ids) inapi-policy.yaml, written byapi approval NAME --decide-with relayed --relayers concierge(a loosening, reviewed like new approvers;--decide-with directnarrows it
again). The default staysdirect: the person decides with their own credentials. A relayed
decision is accepted only from a listed agent, on a thread it started for that user, with
requesteran approver, naming the approval'sdigest(a SHA-256 of the call as the
approver saw it; missing or different: 409approval_digest_mismatch), and is recorded as
decided_via. How the approvers decide is bound to the approval when it is asked, as the
approvers are: a policy that starts or stops relaying, or changes the relayers, while a call
waits does not keep its approval. The approval object showsdecide_with,decided_via
anddigest; the HTTP decision body and the A2A decision part take an optionaldigest.
Rules that decide differently are different gates for the rule-conflict check,api show
andlintprintapproved by requester; relayed by concierge, andapi show --jsonadds
decide_withandrelayersto a relayed gate. The approvals table gainsdecide_with,
relayers,decided_viaanddisplay_digestat startup, and thelanggraph devapprovals
file moves to version 2 (a version-1 file is read, its approvals direct). - A relayed decision shows what it decides and what will happen. An approval of an A2A
message that approves another agent's approval carriesnested(that approval, as the
message sends it: its call, reason, expiry, digest, and the approval it relays in turn) and
effect(the call that will actually happen, the agent that makes it and the agentsvia
which), in/chat,GET /approvals, the thread's approvals and the A2A approval request. It
expires 5 s before the approval it decides at the latest, and the nested calls' and the
effect's query and body are dropped on decision with the call's own (unless
TRACE_CAPTURE=full).approvals listandrunprint the effect first ("orders (via
billing) will POST /orders/ORD-1002/cancel (cancelOrder), as reported by orders", its body,
and avialine per agent), terminal-safe like the rest of the approval. With
requester_actoranddecided_viaon every approval, this narrows KI-009 (approvers still
see the requester as a hash). - The A2A server speaks to agents calling for a user. An agent's card declares the
graph-agents-cli origin extension
(https://ss7172.github.io/graph-agents-cli/a2a/ext/origin/v1, optional): an agent calling
for a user may put the user's own words in the message metadata under that URI (origin:
text,truncated,hops), and for a delegated caller only they reach the run's private
credentials (@origin, whererequire_user_mentionedand the model's note read them),
capped atA2A_ORIGIN_MAX_CHARS(4000). They are never stored: every task is saved without
them. The run a decision resumes acts on the words of the request that paused it, whatever
words the decision carries (the person's "yes, go ahead" at the agent that relays it, or
none): the approval keeps them while it waits (never shown, and dropped once it is decided or
expired, whateverTRACE_CAPTUREsays; fastapi runtime only, as LangGraph Server passes no
credentials to tools), sorequire_user_mentionedholds again on the resumed run and an
agent relaying the approval one level further rebuilds the very call the person approved.
MorehopsthanAUTH_MAX_DELEGATION_DEPTHfails the task (delegation chain too deep). A
task waiting for approval carriesapproval_json, the approvals as exact JSON text, and its
text shows each call's body (up to 2,000 characters) (KI-026). A failed task, or a refused
decision, carries a data part{"type": "error", "code": ...}(thread_busy,
approval_direct_only, ...). A decision may be sent on the context alone, naming the waiting
task inreferenceTaskIds. - A2A tasks follow their approval's outcome (KI-025). However an approval ends (a decision
over A2A, on the task or on its context; one taken over HTTP, the person at this agent
included; or its expiry), the requester'sinput-requiredtasks on that thread that wait on
it take the resumed run's outcome (completed,failed, orinput-requiredwith the new
approvals) and say where it continued (Continued in task <id>.,... was approved outside this task; the run continued there.,... expired before anyone decided.), with the run's
reply. Both task stores; another principal's task on the thread is left alone.role:gates
are decided over HTTP, not A2A: documented as a design choice. - Agents call other agents for the user with a token exchanged for theirs (
auth: exchange,
RFC 8693). An API declaredauth: exchangewithexchange: {audience, scope, resource}is
called withBearer <token>(inforward_header, defaultAuthorization): a token the
identity provider mints for that audience in exchange for the caller's own verified token,
naming this agent as the actor. It is asked for just before the call is sent
(app_utils/token_exchange.py), after the policy check, the approval gate and the limits: a
refused call, or one paused for a person's approval, never exchanges, and nothing is
exchanged while a request is authenticated. Tokens are kept in process memory only, per user
token, audience, scope and resource, for at mostexpires_in, 300 s
(TOKEN_EXCHANGE_MAX_TTL_S) and the user's own token's expiry, less 30 s; concurrent calls
share one exchange. A user token with 10 s or less left is not exchanged; an issuer refusal
is remembered forTOKEN_EXCHANGE_FAILURE_TTL_S(10 s); three issuer failures in a row
(timeouts afterTOKEN_EXCHANGE_TIMEOUT_MS, 2 s; connection errors; 5xx; unusable answers)
open a circuit breaker that fails calls to exchange APIs at once, then lets one call probe
the issuer. Nothing is sent when there is no token to exchange (ashared-bearercaller, a
run resumed by a role approver), and the tool reads why. The issuer and this agent's client
areTOKEN_EXCHANGE_URL(https outside dev unless loopback or
TOKEN_EXCHANGE_ALLOW_HTTP=true),TOKEN_EXCHANGE_CLIENT_ID,TOKEN_EXCHANGE_CLIENT_SECRET
(client_secret_basic, orclient_secret_postthroughTOKEN_EXCHANGE_CLIENT_AUTH);
outsideAPP_ENV=devthe app refuses to start with an exchange API and any of the three
missing, or undershared-bearerorlanggraph-server, and a malformed setting stops it
everywhere. Exchange APIs receive the request'sX-Request-IDand trace context, asauth: forwardones do (the owner's decision). A call to an agent already in the request's
delegation chain, or to this agent itself, is refused before anything is sent. New metrics:
agent_token_exchanges_total{api, outcome}(issued,cached,refused,no_actor,
unavailable,circuit_open) andagent_token_exchange_duration_seconds{api}; one log line
per exchange sent and a warning when the breaker opens, never a token. An exchanged token
that names no actor is refused, and nothing is sent (the owner's decision of 2026-09-28): a
JWT without theactclaim (or the oneAUTH_JWT_ACTOR_CLAIMnames; some issuers, Keycloak
among them, add none), or a token the agent cannot read as a JWT (opaque, encrypted). The
called agent would read such a token as the user's own, and could let this agent decide the
user's approvals there. The refusal is remembered forTOKEN_EXCHANGE_FAILURE_TTL_S(metric
outcomeno_actor). An API opts in withexchange.allow_actorless: true, for a called agent
that setsAUTH_JWT_DIRECT_CLIENTSand lists this agent asclient:<its client id>in
AUTH_ALLOWED_ACTORS; the first such token then logs a warning (KI-149).jwtnow also
keeps the token'sexpfor this (keep_subject_tokentakesexp). The authentication guide
has a Keycloak recipe. Existing projects get the runtime fromscaffold upgrade(a new
app_utils/token_exchange.py;api_client.py,auth.py,metrics.py,telemetry.py,
fast_api_app.py). auth: forwardcan forward the caller's own token to an agent it was minted for:
forward_audience. Underjwt, where no per-API credential is set, anauth: forwardAPI
withforward_audiencesendsBearer <the caller's token>only when that token'saud
names the audience; a token minted for this agent alone is never replayed at another. Prefer
auth: exchange.api add --auth exchange --audience AUD [--scope S] [--resource URI] [--allow-actorless]
declares an exchange API (--allow-actorlesswritesexchange.allow_actorless: true);
--audiencewith--auth forwardwritesforward_audience, and--forward-headergoes
with either mode. The first exchange API addsTOKEN_EXCHANGE_CLIENT_SECRETto the
manifest'ssecrets.keys,TOKEN_EXCHANGE_URLandTOKEN_EXCHANGE_CLIENT_IDto
.env.example(the secret commented out) and to the chart'svalues.yaml(a placeholder URL
and the project's name), and lists what is left: the issuer's permission to exchange for the
audience, and the callee'sAUTH_JWT_AUDIENCEandAUTH_ALLOWED_ACTORS; without the opt-in,
that the issuer must name this agent inact(else calls fail), and what the opt-in needs;
with it, the callee'sAUTH_JWT_DIRECT_CLIENTSand this agent asclient:<id>in
AUTH_ALLOWED_ACTORS.lintnotes the same for every API that opts in.api removetakes
them away with the last exchange API, andcreate --api-policywith an exchange API renders
the same.api shownames the audience, scope and resource, and the opt-in. The schema
acceptsforward_headerwithexchangeas well asforward(both SHARED copies; the
messages name every mode), andcreate,lintandapi addrefuseauth: exchangeunder
langgraph-serveras they refuseauth: forward.- JSON-RPC APIs and other agents in
api-policy.yaml:protocol,rpc_methodand
a2a_operation. An API may setprotocol: jsonrpc(a JSON-RPC 2.0 API) orprotocol: a2a
(another agent over A2A 1.0 JSON-RPC, with its endpoint ina2a: {path: /a2a/<name>}), and
any API adescription. For those, every POST must send one JSON-RPC request object (a
batch, a notification, extra members, a body that is not plain JSON, or a body on a GET or
HEAD is refused before sending), and the policy client reads what it is from the body, never
from the tool's label: its method (rpc_method; undera2aan A2A 0.3 name such as
tasks/cancelis read as its 1.0 name,CancelTask) and, for an A2A message whose parts
name an approval, what it decides (a2a_operation:rejectonly when every such part
rejects, elseapprove; a message method in any letter case is read for a decision).
Operation entries may pinrpc_methodanda2a_operation: an allow must match them; a
denial or gate covers every call they describe whatever its path, label or spelling. A
protocol: a2aAPI that can send messages must gate or denya2a_operation: approve, so an
agent never decides on its own an approval the agent it calls waits for (the policy is
invalid otherwise, and the client refuses such a message at runtime too), and refusesauth: none; JSON-RPC APIs allow GET, POST and HEAD only. A tool'soperation_idthat names an
entry pinning another method or decision is refused (the label cannot hide a request). Plain
httpAPIs, and every entry without the new keys, are judged exactly as in 0.2 (property
tests against the 0.2.0 rules). Withprotocol: a2a, the transport rule for peers applies:
outsideAPP_ENV=deva credential goes to such an API over https only, unless the host is
loopback, a single-label or a.svcname (<ENV> must use https outside APP_ENV=dev to carry credentials); other APIs keep their transport. Both SHARED copies; every message is in the
schema reference. KI-122 is closed for these APIs: their gates namerpc_methodor
a2a_operation, which no label can hide. lintandapifor JSON-RPC APIs and A2A peers.API_CALLSentries takerpc_method
anda2a_operation: every POST to aprotocol: jsonrpc|a2aAPI declares itsrpc_method
(an A2A 0.3 name is read as its 1.0 name), and lint judges the declaration as the client
judges the body, refusing the keys on anhttpAPI or a GET, anoperation_idthat names
another request, and a message that approves with nothing holding it; a refused JSON-RPC call
gets itsapi allow ... --rpc-method Mhint.lintwarns about an A2A peer without a
descriptionand about more than 40 peers. New flags:api add --protocol http|jsonrpc|a2a --a2a-path P --description T(ana2aAPI that allows POST is written with
denied_operations: [{a2a_operation: approve}], fail closed, and the note says how to gate
the person's decision instead;--auth noneis refused),api allow NAME [OPID] --rpc-method M --method POST --path P,api deny NAME --rpc-method M/--a2a-operation approve|reject,
api revokewith the same two, andapi approval NAME --a2a-operations approve[,reject]|none
(which keeps the rule's--operationsentries, and the other way round).api showprints
the description and the protocol, and--jsonaddsdescription,protocolanda2aper
API andrpc_methodanda2a_operationper declared call.- A decision is bound to the JSON-RPC request it was taken for. On a
protocol: jsonrpc|a2aAPI every call is a POST to one endpoint, so API, method and path could not
tell a relay's calls apart: the read a resumed tool sends first would have taken the
decision its approve message waits for, and a rejection would have stopped thereject
message that tells the other agent. The call's identity now includes its JSON-RPC method and
A2A decision (read from the body) for such APIs, in the interrupt, the approval record (its
payload), the decision and the ledger's bound approvals; the approval object and its
digestinclude them. Calls tohttpAPIs keep their three-field identity, and their
records, decisions and digests are unchanged. - Requests keep one request id and one trace across agents (
PROPAGATE_TRACE_HEADERS).
An agent that called another agent started a new request id and a new trace there, so a
request across agents could not be followed from end to end. Under
PROPAGATE_TRACE_HEADERS=peers(the default), every call through the policy client to
another agent (protocol: a2a, whatever itsauth) or to an API that acts for the calling
user (auth: forward,auth: exchange) carries the request'sX-Request-IDand, under
OTLP tracing, its W3C trace context (traceparent,tracestate). Other APIs (auth: bearerorauth: noneoverhttporjsonrpc) are third parties and never receive them.
An incomingtraceparentis continued on the A2A routes (/a2a/*) only, so a caller of
the public routes (/chat, the thread and approval routes) cannot choose the agent's trace
ids; itsX-Request-IDis still taken and echoed.allsends them to every API and
continues a trace on every path (an agent behind a tracing gateway);offneither.true
reads aspeers(logged once) andfalse(also0,no) asoff; any other value stops
startup. Undershared-bearerandlanggraph-server, which refuseforwardand
exchange, only an agent's peers receive them. The headers are not bound by approvals.
Existing projects get it fromscaffold upgrade(app_utils/telemetry.py,
middleware.py,api_client.py,fast_api_app.py). - An A2A client in the template:
app_utils/a2a_client.py(B3; replaces the
experiment's hand-written peers client).peer_tools(PEERS)gives the model
ask_agent(agent, request)(its description lists the peers and what each does) and, for
peers it relays approvals to,approve_agent_action(agent, task_id);A2APeerClientdoes
the same from your own tools (send,get_task,pending_approvals,decide,cancel,
relay,card). Every request goes through the API policy (the allow-list, the approve
gate,auth: exchange, the limits, the response cap); the SDK's HTTP client is never used.
Before the first call a peer's agent card must offer an A2A 1.x JSON-RPC interface at the
URL this agent calls, named aftera2a.path(cachedA2A_CARD_TTL_S, 300 s; its URL never
dialed). ThecontextIdis a UUID keyed withPRINCIPAL_HASH_SALT, one per thread, peer and
user; calls to one peer in a thread are serialized; a busy peer is asked again 3 times, and a
cancel another replica runs once. Replies are the lastresponseartifact (at most
A2A_REPLY_MAX_CHARS, 6000). WithA2A_FORWARD_ORIGIN=auto(default) a peer whose card
declares the origin extension gets the user's own words. Calls back to this agent or up the
delegation chain are refused. The relay reads what the peer waits on from the peer (its
exactapproval_json, or its approvals ledger when it lost the task), reports a gate the
person decides at the peer asneeds_direct_approval, and otherwise sends one
context-addressed decision, the same on every run, through this agent's approve gate: the
person approves it here, seeing the peer's call aseffect; a rejection tells the peer at
once..env.exampleand the environment reference list the new settings, which the startup
check validates. The fake test model fills a one-valueLiteralargument (a JSON-schema
const). graph-agents-cli peer add|remove|list|show|sync: declare the agents this agent asks
(B3).peer add NAMEwrites the peer'sprotocol: a2aAPI (its endpoint/a2a/NAME, the
credential the auth policy calls for:jwtexchange for audience NAME,customforward,
shared-bearerbearer; the card,SendMessage,GetTask, optionallyCancelTask; the approve
gate with--approvals relay, the default, or its denial withdeny; 12 calls a run, a 120 s
read timeout and a 1 MiB answer cap), the manifest'ssecrets.keys,.env.example, the
chart'svalues.yamland, with--cluster-url, eachvalues-<env>.yaml(never.env), and
regenerates<agent_dir>/tools/a2a_peers.py: data only (PEERS, a literalAPI_CALLSof
exactly the calls the policy allows,TOOLS = peer_tools(PEERS)), importing the manifest's
agent directory. It prints what is left, including the peer's side:AUTH_JWT_AUDIENCEand
AUTH_ALLOWED_ACTORSthere,AUTH_JWT_DIRECT_CLIENTSandclient:<id>for an issuer whose
exchanged tokens name no actor (--allow-actorless), and theapi approval <its gated API> --decide-with relayed --relayers <this agent>line that lets it relay (a reviewed loosening
the peer's owners run). Guards: a 0.2 runtime (runscaffold upgrade), a bad or own name
(exit 2), a taken API name, a peer with other settings, an auth mode the project cannot serve,
a module it did not write (exit 3);--card URL|FILEreads the description, path and origin
support.peer removetakes it all back,peer listandpeer show [--check]report (the
check reads the card without a credential: reachable, 401, or a foreign endpoint, exit 1),
andpeer syncregenerates the module afterscaffold upgradeorapiedits.lintfails
while the module and the policy differ, and notes a tool module that calls a peer itself.
Wiring the round-1 concierge's five peers takes fivepeer addcommands instead of 40api
commands and 418 hand-written lines.deploychecks the agents the project calls before building (the peer pre-checks).
Outside dev (the environment or the pods'APP_ENV) it refuses, exit 3, in every CD mode:
a credential for aprotocol: a2aAPI sent over plain http to a host that is not loopback, a
single-label name or a.svcname (the runtime's transport rule); anauth: exchangeAPI
withoutTOKEN_EXCHANGE_URLorTOKEN_EXCHANGE_CLIENT_IDin the chartenv; a plain-http
TOKEN_EXCHANGE_URLto a host that is not loopback withoutTOKEN_EXCHANGE_ALLOW_HTTP. It
warns about an unset peer URL and aboutTOKEN_EXCHANGE_CLIENT_SECRETor
PRINCIPAL_HASH_SALTmissing fromsecrets.keys; in dev the refusals warn too. A test keeps
the CLI's copy of the runtime'sinternal_hostequal to the template's.graph-agents-cli system check|apply|graph|delegations|deploy: agent projects that call
each other, seen as one (B13). An optionalgraph-agents-system.yaml(found upward, or
--file; JSON Schemaschemas/graph-agents-system.schema.json, generated from the models)
names each agent's project, its client id and actor id (actor_id: theact.subof its
exchanged tokens, which the agents it calls list and--relayersnames; default the client
id, for an issuer that names the client otherwise there, such asagent:<client>, where
listing the client id would leave every call refused with 403 whilesystem checkpassed:
found by the round-3 acceptance run), the agents it calls (withapprovals,auth,
scope,calls,description), the environments (port_basefor local processes, aurl
template, or in-cluster URLs from each chart and manifest namespace), the token issuer, a
shared database'smax_connectionsanddeploy.parallel. A file that cannot be used exits 3
(an unknown project or one two agents name, an edge to an unknown agent or to itself, one
agent called twice by another, two agents with one client id or actor id, anexchangeedge without
identity, an environment a manifest does not know).system applywrites both sides of
every edge, idempotently and one diff per project: in each caller whatpeer addwrites, and
per environment the peer URLs,TOKEN_EXCHANGE_URLandnetworkPolicy.egressToto the
called agent's pods (TOKEN_EXCHANGE_CLIENT_IDis the file's client id); in each called
agent itsappUrlper environment,AUTH_JWT_AUDIENCEwhen empty, the callers' actor ids
inAUTH_ALLOWED_ACTORS, andnetworkPolicy.ingressFromfor the callers' pods (plus the
Gateway's namespace while the route publishes paths). It never writes gates (it prints the
api approval ... --decide-with relayedline a relay needs), secrets,.envor local
settings, removes a peer of the file's agents that leftcalls, keeps an existing peer's
tuning, and never takes access away.system checkruns SC01-SC13 (runtimes and charts,
edges in step, paths, issuer and audience, appUrl against the URL dialled, replicas with
in-memory tasks, relays the called agent's gates refuse, allowed actors, cycles and
delegation depth, a shared database's connection budget, the callers' secrets, exchange under
langgraph-server, an A2A path still public), and with--liveSC14-SC15 (Services with a
ready endpoint, read from their EndpointSlices since the v1 Endpoints API is deprecated;
cards and the token URL answering; the Secrets' key names), using each project's recorded
kube context and never the current one outside dev; exit 1 on an error,--json.system graphdraws the system (mermaid, dot or json);system delegationsprints what the issuer
must allow each client and guarantee.system deploy --env ENVchecks, then runs
graph-agents-cli deployin every project in waves, callees first (a cycle broken in file
order), at most--parallelat once (default 3), stopping after a failed wave unless
--keep-going, printing each agent's build, load and rollout times, then checks--live;
outside dev it requires every project's recorded context, and projects inargocdmode (a
commit and a pull request each) go one at a time.api/_filesedits now build on the planned
text of a file, so one plan can hold several peers. The NetworkPolicy rules select an agent's
pods by the chart's selector labels (name and release: the bundled database shares the
release label), and were checked on a kind cluster with kindnet enforcing them: callers reach
a called agent's Service port 80 through an egress rule on the pods' port 8000 (a rule on
port 80 blocks them), and a pod in another namespace is refused. KI-159 and KI-160 park the
residuals (LangGraph Server's own pool in SC10, a local environment's.envunchecked);
KI-158 (SC14 on the deprecated Endpoints API) was fixed before release.- A guide to agents calling agents
(website/src/guides/multi-agent.md): who acts for whom,
a walk-through frompeer addto an eval at the entry agent, relayed approvals, the user's
own words, the system view, following one request across agents, sizing, the threat model
with what remains, and the limits. The security checklist adds the internal A2A paths, what
the issuer must allow and the salt for agents that ask others; the observability guide names
theactorlog field; the manifest reference says who addsTOKEN_EXCHANGE_CLIENT_SECRET
and peers' keys tosecrets.keys. A reference page,
graph-agents-system.yaml, lists every key of the
system file with its default and what makes a file unusable. - The skills cover agents calling agents. The workflow skill asks in Phase 0 whether the
agent asks other agents or is called by them, declares peers withpeer addorsystem apply(never a hand-written client), runslint,peer show --checkorsystem check,
evaluates at the entry agent and deploys withsystem deploy. The langgraph-code skill has
a section on the generatedtools/a2a_peers.pyandA2APeerClient, what tools see when an
agent asks for the user (current_caller,require_direct_caller,require_owner,
require_user_mentionedunderA2A_DELEGATED_MENTIONS) and relays (decide_with). The
deploy skill coverssystem apply|check|deploy|delegations,AUTH_ALLOWED_ACTORS,appUrl
per environment, keeping A2A paths internal and sizing a shared database; the scaffold skill
sayspeer addwrites the client side and thatlanggraph-servercalls other agents only
withauth: bearer. Each changed skill passes gac-bench's fact-check (every command and
option exists; within 1.25 times the 0.2.0 text). limits.max_response_bytescaps an API's answers. With it (1 to 67108864 bytes), the
client reads a response body, decoded, only up to that many bytes: past it, the answer is
discarded and the call fails with<api> answered with more than N bytes; discarded(a
declaredContent-Lengthover the cap is refused before reading). Memory stays bounded
whatever the answer's compression: a capped call asks for gzip or deflate at most and
decodes the body itself, never past the cap (httpx decodes each network read whole, and 32
KiB of zstd is 1 GiB), and an answer in any other content encoding (zstd, br, an unknown or
a stacked one) is refused unread:<api> answered in a content encoding other than gzip or deflate, which limits.max_response_bytes cannot bound; discarded.api add --max-response-bytes Nandapi limits NAME --max-response-bytes N|noneset it, andapi limitsedits keep it;api showprints it. Unset, answers are read whole as in 0.2 (no
default cap for existing APIs). Both SHARED copies.- Reasoning effort and the Responses API for OpenAI-API models:
MODEL_REASONING_EFFORT
(none,minimal,low,medium,high,xhigh) andMODEL_USE_RESPONSES_API(true:
every request to/v1/responses;false: Chat Completions; unset: langchain-openai
chooses, as before), foropenaiandopenai-compatible; the judge reads
JUDGE_REASONING_EFFORTandJUDGE_USE_RESPONSES_API, defaulting to the agent's when it is
an OpenAI-API model too. A model that refuses function tools with a reasoning effort on Chat
Completions (the experiments' finding F15: its first call failed with a 400 naming
/v1/responses) now works with the switch on, tools, streaming and token usage included. A
bad value, or either setting for another provider, stops startup..env.exampleand, for
OpenAI-API projects, the chart'svalues.yamldocument them; existing projects get the
runtime fromscaffold upgrade(app_utils/model.py,fast_api_app.py). - Structured final answers: a JSON schema for the agent's answer. A project that puts a
JSON Schema (root"type": "object") in<agent directory>/response_schema.json
(create --response-schema FILEseeds it) has its agent answer in that shape.agent.py
builds the agent withresponse_format=response_format(model, tools)
(app_utils/structured.py), one of LangChain'screate_agentstrategies:
RESPONSE_FORMAT_STRATEGY=auto(default) uses the provider's own structured output where
LangChain's model profile says the model has it with the agent's tools bound and the
model's client can send the schema, strict on OpenAI (LangChain's own auto mode asks OpenAI
for a best-effort schema), else afinal_answertool the model must call (tool_choice
forces a call at every step);providerandtoolforce one. Anthropic's client refuses a
type list and a schema with notype(anenumalone) before any request, soautouses
the tool for such a schema andproviderstops startup naming why. LangChain returns a raw
JSON-schema answer unchecked, so the newStructuredAnswermiddleware (last in
middleware()) checks every answer against the schema: one that does not fit, a reply that
is not JSON, a plain-text final reply, or an answer given beside other tool calls (none of
which runs, so a gated call never runs after an answer already given) goes back to the model
with what is wrong, up to 3 tries in the same step (the failed tries stay out of the thread;
their tokens count in the answer's usage), then the run ends with the newerrorcode
invalid_structured_response. The checker supports a documented JSON
Schema subset and refuses a schema that uses anything else at startup (lintand
create --response-schemaapply the same rules, a SHARED block kept byte-identical with the
template's). Delivery: a completed/chatrun's onlymessage.deltais the answer's JSON
text andmessage.endcarries the object asstructured_response; the answer tool never
shows as atool.call; a run paused for an approval answers once resumed; the A2A
responseartifact adds a data part with the object (mediaTypeapplication/json,
streamed or not; an A2A task that waited on an approval decided elsewhere, over HTTP
say, takes the whole answer as its lastresponseartifact) and the card lists
application/jsonamong its output modes;eval
recordsstructured_responsein its traces andexpect.json_schemachecks it as it is.
Both runtimes, both checkpointers. Without the file nothing changes. Measured with
gpt-5-mini on a 24-case triage task (a tool call, then a six-field answer), twice: without
the mode 28 of 48 replies were not a bare JSON document (the sentence the model writes
before a tool call came first), with it 0 of 96 (both strategies, no correction needed),
at the same cost for the provider strategy and about twice the output tokens for the tool
strategy. Existing projects:
scaffold upgradebringsstructured.pyand the runtime, but never rewritesagent.py:
passresponse_format=response_format(model, tools)tocreate_agentand add
StructuredAnswer()last tomiddleware()before adding a schema (lintwarns about
either). The runtime checks every answer again before delivering it, so an answer that
never went throughStructuredAnswerand does not fit ends the run with
invalid_structured_responseand is not sent on/chator over A2A (it stays in the
thread, KI-172). Project tests: a generated project's tests run with the mode off
(RESPONSE_SCHEMA_PATH=none, a new value: no schema, whatever the file), so a project with
a schema keeps a green suite and CI, andtests/unit/test_structured.pychecks the
project's own schema and that one turn ofagent.pyanswers in it; the fake model takes a
choice of the schema that fits (nullfor the develop guide's optionalorder_id) where
its text breaks apattern. Documented in
Develop your agent, the HTTP API
and environment references, the upgrading guide (0.2 to 0.3) and the multi-agent guide (a
peer that answers in JSON); the langgraph-code, workflow, scaffold and eval skills teach it
(declare the shape, never parse JSON out of a reply, wire a 0.2agent.pyby hand); and
gac-bench has a task family for it. KI-165 to KI-177 park its remaining minor issues
(KI-168 also covers A2A). - A documentation site in
website/(MkDocs Material): Get started (installation, a
five-minute quickstart, two tutorials, the lifecycle), guides for building and operating an
agent, and a reference whose CLI and Skills pages are generated from the commands and
skills/..github/workflows/docs.ymlbuilds it withmkdocs build --strictand checks its
links on every pull request, and publishes it to GitHub Pages frommainwhen the
repository variablePUBLISH_DOCSistrue: the site is at
https://ss7172.github.io/graph-agents-cli/. Preview it with
uv run --group docs mkdocs serve -f website/mkdocs.yml. - gac-bench, contributor tooling for the skills (
tools/skillopt/): a benchmark of
realistic graph-agents-cli tasks for each of the six skills, 104 of them, including this
release's agent-to-agent features (peer add, relayed approval gates,auth: exchange,
rpc_methodrules andsystem apply) and structured final answers, each with a
deterministic verifier and scripted gold and broken solutions, in frozen train, val and
test splits; and
gac_skillopt, an environment in which SkillOpt
runs a candidate skill in Claude Code or Codex, isolated so that the session sees only that
skill, scores it and proposes edits. It found the skill rules this release adopts after
review, andtools/skillopt/results/keeps every measurement. It is never a dependency of
the CLI or of a generated project and is in neither the wheel nor the sdist (a fast test
guards the build configuration); CI runs its unit tests. CONTRIBUTING.md and the site's
Skills benchmark page say how to run it.
Changed
- The README is a short entry point with absolute links (it is also the PyPI page). Its
former sections, including "Known limitations" and "Where it is behind", moved to the
site's pages: each limitation now sits on the page of the feature it concerns, the
comparison on Compared with google-agents-cli. The
tests that ran README examples now run the same examples from the site's guides. - The workflow and scaffold skills carry rules found by the SkillOpt experiment.
SkillOpt optimised the skills against gac-bench, a
benchmark of graph-agents-cli tasks carried out by Claude Code and Codex sessions. Its
proposals were reviewed by hand (wording taken from the benchmark generalised, one claim
corrected, a misplaced rule moved, a rule learned from the benchmark's harness dropped) and
measured again before they were adopted. The workflow skill now scopes Phase 0 and its spec
gate to a new agent: a concrete change to an existing project (add a retry to a tool, bump
a dependency, fix a crash) is not a new agent, and is made without a new spec, while the
decisions it raises that the user owns (new API operations or access, approval gates, the
model) still go to the user. This rule was written by hand after an earlier candidate made
Codex refuse such changes. For a new agent it says what counts as approval of the spec: a
request to build, a draft spec, defaults the agent chose or an instruction to proceed on its
own is not one. With nobody to approve, the agent stops beforecreateat a draft spec and
ends its answer with each open decision as a direct question, and asks for approval as a
question. It also says to fix the agent, not the eval, when an eval that passed breaks
(unless the user asked for the change the eval checks), maps failed checks to code, says how
to show that a failing test is unrelated to a change, proves a change withlint, theeval runcounts and arunsmoke test, and never presents thefakemodel's exit 0 as evidence
of quality. On the benchmark's workflow val tasks, against the 0.2.0 text: with Codex
(gpt-5.6-terra, 2 repetitions) 12 of 12 passed against 6 of 12, the open decisions were
asked as questions in 6 of 6 spec-gate sessions against 0 of 6, and every requested change
was still made (6 of 6); with Claude Code (3 repetitions) 17 of 18 passed against 10 of 18,
the spec-gate sessions stopped beforecreatein 9 of 9 against 1 of 9 and asked in 9 of 9
against 1 of 9. The adopted wording, with generic examples, was checked again on Claude
Code: 18 of 18, with changes made, stops and questions each 9 of 9. The scaffold skill keeps
the defaultAGENTS.mdguidance file unless the user says the team uses one coding agent
(8 of 8 sessions, against 4 of 8), its examples no longer pass
--agent-guidance-filename CLAUDE.md, and it has the agent readcreate_paramsbefore
scaffold enhanceand aftercreate. - The observability skill has a procedure for salting the hashed principal id, found by
the SkillOpt experiment with Codex sessions and reviewed by hand before it was adopted: add
PRINCIPAL_HASH_SALTtosecrets.keys(configuration only: no code, chart template or
values change), put the value in.env.<env>, thensecrets apply --env <env>and
deploy --restart --env <env>for every deployed environment, and name those commands
verbatim when the user runs them. Codex followed it in 6 of 6 sessions of the benchmark's
salt task, against 1 of 9 with the old text. Tracing for one environment now edits only
that environment's files and ends with the commands still to run; local tracing edits
only.envand adds none of the LangSmith SDK's own switches. Its paragraph on trace
headers across agents states this release's rule (PROPAGATE_TRACE_HEADERS=peers|all|off:
other agents andauth: forward/exchangeAPIs, and an incoming trace continued on
/a2a/*only) instead of 0.2's. - Approval times read from Postgres are in UTC, whatever the database session's time zone
(created_at,expires_at,decided_at,used_at; they came back in that zone, the same
instant written differently). An approval now reads back with the very times it was
created with, which the A2A client's relay binds in its decision when it reads what a peer
waits on from the peer's approvals ledger (the peer lost the task): on a database not set
to UTC that decision differed from the one the person approved, and nothing was sent.
Fixed
eval'sjson_schemacheck reads the reply's answer, not its first JSON. It parsed the
first fenced code block, and otherwise everything from the first{or[to the end of
the reply: a reply that showed an example (or quoted its input) before its answer was
checked against the example, and one that added prose after raw JSON failed as "not valid
JSON" (the experiments' finding F10). It now reads the whole reply when that is JSON, else
the reply's last JSON object or array, taking the schema's root type when it names
objectorarray(so a citation such as[1]after an object answer is skipped), inside
code fences or not.run,eval run,eval generateandapprovalsno longer fail under a SOCKS proxy.
WithALL_PROXY=socks5h://...in the environment (as coding-agent sandboxes such as Codex's
network proxy set it), every request crashed withImportError: Using SOCKS proxy, but the 'socksio' package is not installed, even to the command's own local server and even when
NO_PROXYlisted it: httpx builds a transport for every proxy variable when a client is
created. Requests to this machine (the local server, the playground) now never use the
environment's proxies. Requests to another machine still honourHTTP_PROXY,HTTPS_PROXY
andNO_PROXY, and SOCKS proxies now work too: graph-agents-cli depends onhttpx[socks]
(adds the pure-Pythonsocksio). A proxy setting httpx cannot use (another scheme) is a
one-line error naming the variable, exit 3, instead of a traceback;eval generate --url
reports it once, before any case runs.- The local server is stopped even where listing processes is denied. In a sandbox that
refuses the process table (Codex's seatbelt denieskern.proc.all), psutil raised
PermissionErrorwhilerun,eval runoreval generatestopped the server they
started: the command ended withError: PermissionError: [Errno 1] Operation not permittedand the server kept its port, orphaned. The teardown now signals the server's
own process group (it is started in a new session, and its uvicorn child stays in it) and,
for a server this invocation started, its process handle; no error from the teardown
replaces the error of a failed start any more. run --stop-serverno longer reports success for a server it could not stop. When the
operating system refused the signal (Claude Code's sandbox lets a command signal only the
processes it started itself, so a server started by an earlier command survives), it
printed "Local server stopped." and deleted the record while the server kept its port; the
nexteval runthen failed with "something is already listening". It now exits 2, names
the processes still running and how to stop them, and keeps the record, so later runs reuse
that server. A server thatrun,approvalsoreval generatecannot stop after their
work is a warning that never replaces their own result or error. Neither does a server a
command would replace (idle for 30 minutes, or no longer answering) but may not stop: a
warning names it and thekillcommand, then the command reuses it while it still answers,
or starts a fresh one beside it.infohonoursGRAPH_AGENTS_CLI_NO_UPDATE_CHECK=1and CI. It rannpx -y skills@1.5.9 list --jsonevery time, which may download the package: a disconnected
install waited up to 15 s for it, and in a coding agent's sandbox the blocked download read
as "the CLI is blocked". With the variable set (the disconnected profile) or a CI marker, the
listing is skipped like the skills version check, andinfosays so ("Installed skills:
not listed (... skipsnpx skills list)";--jsonaddsinstalled_skills_skippedwith the
reason, null when the listing ran).eval gradeworks in a project with another agent directory. The judge runner
importedapp.app_utils.modelwhatever the project's package was, so in a project created
withcreate --agent-directory <name>every judge metric (the default dataset's
response_qualityincluded) failed with exit 3, "the judge model is unreachable or
misconfigured". The runner now imports the judge from the manifest'sagent_directory.deployworks for a project that needs no Secret key. An env file that sets none of
the allow-listed keys (a keyless project: thefakemodel, a keylessopenai-compatible
endpoint, noshared-bearerkey) madedeployexit 3, "No allow-listed secret values to
apply", and only after it had built the image and loaded or pushed it. Such an env file now
leaves the Secret as it is, like a missing env file indev, and says so; the decision is
made before anything is built. Outsidedev, where the chart requires the Secret
(secretOptional: false), a missing Secret stops the deploy (exit 1) before the build.
secrets applystill exits 3 when there is nothing to apply.- A2A tasks are shared by every replica and survive restarts (KI-024, now Low). The task
store was in process memory, per replica, while the production values run two replicas:
GetTask,ListTasks,CancelTaskand a message naming ataskId(an approval decision)
failed with -32001 "Task not found" whenever a request reached the other pod (7 of 30
GetTaskcalls over new connections in the A2A experiment), and a restart or rollout
dropped every task, including those waiting for an approval. UnderCHECKPOINTER=postgres
(and a PostgresDATABASE_URIunderlanggraph-server) tasks are now kept in the app's
database, tablea2a_tasks(agent_a2a_tasks), with the same per-principal ownership,
A2A_TASK_TTL_Sexpiry (on the database clock;0now keeps a task until its thread is
deleted) and deletion with their thread, on every replica. A task whose run ended with its
process turnsfailedinstead of stayingworking. ASubscribeToTaskorCancelTask
that reaches a replica other than the one running the task is refused (-32004, -32002)
instead of waiting for events that happen elsewhere, or reporting a cancel the run then
overwrites. Streamed reply chunks are written at most once a second per task. Postgres
cannot store the character U+0000, so a stored task holds U+FFFD in its place (a message,
a tool's output or a contextId holding one never fails the task's save).
CHECKPOINTER=memorykeeps the in-memory store. The a2a SDK'sDatabaseTaskStorewas
not used: its Postgres extra is SQLAlchemy on asyncpg (the template uses psycopg only),
itscontext_idcolumn holds 36 characters where the template accepts 128, and it
creates its table on first use, outside the schema lock that replicas share. Existing
projects get it fromscaffold upgrade(it changesapp_utils/a2a.pyand
app_utils/db.py); the table is created at the next start. - A tool's output that is not valid Unicode no longer breaks the run. A lone surrogate
in a tool's result (an upstream JSON"\ud800"escape decodes to one, and UTF-8 cannot
encode it) ended the/chatstream at itstool.resultevent, and failed the A2A task
with-32603and the Python exception text ('utf-8' codec can't encode character ...: surrogates not allowed); underlanggraph-serverthe run failed (run_failed). With a
real model provider the model's next request could not be encoded either, and under
CHECKPOINTER=memoryevery later turn of that thread failed the same way. Each time, the
tool had already acted. The agent middlewareUntrustedToolResultsnow replaces each lone
surrogate with U+FFFD as the result leaves the tool, and the model's request, the/chat
events, the thread history and A2A replies are sent as valid text whatever their source.
Existing projects get it fromscaffold upgrade(it changesapp_utils/content.py,
chat.pyanda2a.py); anagent.pythat builds its own middleware list keeps
UntrustedToolResultsin it. - Streamed runs on an OpenAI-compatible endpoint record their token usage. The generated
agent now asks OpenAI-API models for usage on streamed responses
(stream_options.include_usage); langchain-openai did so by itself only for
api.openai.com, so withMODEL_PROVIDER=openai-compatible(a proxy, a gateway or a
self-hosted server) run records and eval traces showed zero tokens and
expect.max_tokenspassed whatever the run used. On the Responses API
(MODEL_USE_RESPONSES_API) usage comes with every streamed answer anyway, and nothing
more is sent. Existing projects get it fromscaffold upgrade(app_utils/model.py). - KI-146: the observability guide says that under
shared-bearer, as under
langgraph-server, only an agent's peers (protocol: a2a) receive the request id and
trace context. - KI-111: the policy lifecycle no longer starts from a read-only example; the
Outbound API policy guide makes the access level an
explicit choice at every step. - KI-113: the documentation site is published at
ss7172.github.io/graph-agents-cli, so the
README's links to it work.
Security
- Under
langgraph-server, a native run no longer chooses who its tools act for. The
server's native run API (POST /threads/{thread_id}/runs,/runs/wait,/runs/stream
and the thread-less/runsroutes; the default/threadspath prefix publishes the first
three) takes the run context from the request (context, orconfig.configurable, which
the server copies into it), and tools read the calling principal from that context
(current_caller(),require_owner,require_direct_caller,require_user_mentioned).
The auth handler checked whose thread it was but passed the context through, as in 0.2.0:
any authenticated user could have the tools act as another principal, with roles of their
choosing ({"principal_id": "bob", "roles": ["ops"]}), and an agent presenting a user's
request could drop its@actorand passrequire_direct_calleras the user. The handler
now puts the caller's own run context on every run it authorizes, whatever the request
sent: the caller's id, its roles (an agent's are only thoseAUTH_DELEGATED_ROLESlends)
and its public attributes with@actor(never credentials), the context/chatsends
(authenticatepublishes it in the server's user asrun_context). A native run that
sends none now acts for its caller, where every tool that needs one refused (KI-020 no
longer lists that). Studio underlanggraph devkeeps the context it sends. Existing
projects get it fromscaffold upgrade(it changesapp_utils/auth.pyandchat.py). - Agents that call each other for a user fail closed. A request another agent presents
for a user is refused until the called agent lists that agent (AUTH_ALLOWED_ACTORS,
empty by default), keeps none of the user's roles (AUTH_DELEGATED_ROLES, empty), reaches
only the threads, tasks and approvals it started for that user, and never decides an
approval unless the gate relays through it by name (decide_with: relayedwith
relayers, naming the approval's digest; the default isdirect). A calling agent
exchanges tokens only after the policy check and the approval gate, refuses an exchanged
token that names no actor unless the API opts in, sends credentials to other agents over
TLS outside dev (loopback, single-label and.svchosts excepted), and refuses to call an
agent already in the request's chain. A message that
approves another agent's approval must be gated or denied by the caller's policy. These
close the design flaws the A2A multi-agent experiment found in a hand-built system (an agent
holding a user's delegated token could decide that user's approvals; any agent in a chain
could read or cancel another's tasks); 0.2.0 itself had no delegated principals.