Skip to content

graph-agents-cli 0.3.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 16:53
· 9 commits to main since this release

0.3.0 lets agents call other agents for the user they serve, over A2A.
Each agent knows the user and the agent in between, a person's approvals stay with that
person, peer add declares the agents one asks, and graph-agents-cli system checks, wires
and deploys several projects as one. It also keeps A2A tasks in Postgres so that replicas
share them, adds structured final answers (an agent that answers in a JSON shape the
project declares), reasoning effort and the Responses API for OpenAI-API models, a
documentation site, skill rules found with the SkillOpt experiment and the benchmark that
measured them, and fixes. What an upgrade from 0.2.0 changes, and the order to do it in, is
in Upgrading projects; the guide to
the new features is Agents calling agents. Parked
medium- and low-priority issues are listed in KNOWN_ISSUES.md.

Breaking changes and migration

Each change below can need an edit in an existing project; the upgrading guide's
0.2 to 0.3 section lists every other
change in behaviour.

  • jwt reads the RFC 8693 act claim. A token carrying it is an agent's for the user,
    refused (403) until AUTH_ALLOWED_ACTORS lists the agent; set AUTH_JWT_ACTOR_CLAIM=
    (empty) to read every token as the user's own, as 0.2 did. A custom policy that returns an
    invalid id (empty, over 256 characters, or with control characters) now fails the request
    with 500 and logs the bug. A custom policy through which other agents forward users'
    credentials must set Principal.actor (policies/ is not upgraded): see the upgrading
    guide.
  • An API whose auth the project's auth policy can never serve stops the app outside
    dev.
    lint and api add check each API's auth against the project's auth policy:
    auth: exchange or auth: forward under shared-bearer, and auth: forward under jwt
    without forward_audience, are errors (exit 3): such an API never had a credential to send,
    so every call to it failed with "the caller has no credential". Outside APP_ENV=dev the
    running app now refuses to start with such an API too, as it does for auth: exchange (the
    owner's decision of 2026-09-28; also auth: forward under the langgraph-server runtime);
    under dev it logs why and starts. Migration: a 0.2 project with such a forward API
    stops starting outside dev after the upgrade until the API gets forward_audience, moves to
    auth: exchange or is removed (scaffold upgrade never rewrites api-policy.yaml, and
    lint names the API). lint and api add also note a forward API with
    forward_audience under jwt ("prefer auth: exchange").

Added

  • Agents calling agents: the caller's identity. A request another agent presents for a
    user is now told apart from the user's own. Principal.id stays the user (the subject);
    the new Principal.actor names the agent presenting the request (its id, the chain of
    agents before it, and its client). jwt reads the RFC 8693 act claim
    (AUTH_JWT_ACTOR_CLAIM, nested act for earlier agents), the client from azp/client_id
    (AUTH_JWT_CLIENT_CLAIM), and with AUTH_JWT_DIRECT_CLIENTS treats a token with no act
    from any other client as that client's; a custom policy sets actor itself
    (actor_from_claims and keep_subject_token are exported for it). Every policy then goes
    through one rule set (finalize_principal): valid ids, at most AUTH_MAX_DELEGATION_DEPTH
    agents (default 3; else 401), only the agents AUTH_ALLOWED_ACTORS lists (default none;
    else 403), and only the roles AUTH_DELEGATED_ROLES lends. Threads, A2A tasks and approvals
    are owned by the subject and the actor: an agent reaches only what it started for that
    user, while the user owns everything done for them: with their own token they read,
    continue and delete the threads their agents started, decide their approvals, and read,
    list and cancel the A2A tasks those agents started for them (the owner's decision of
    2026-09-28; continuing or subscribing to such a task stays with its agent). A delegated
    principal's roles never read across, administer or decide as a role: approver, and it
    never decides an approval (403 approval_direct_only: the person decides with their own
    credentials). The actor reaches
    tools in attributes["@actor"], is recorded with each approval (requester_actor), and is
    logged (actor), kept in run records and named in trace metadata. auth dev-token mints
    such tokens locally (--act, repeatable, and --azp). The database gains
    threads.actor, runs.actor and approvals.requester_actor at startup (ADD COLUMN IF NOT EXISTS); existing rows are direct, and every 0.2 principal is direct, so 0.2 behaviour is
    unchanged for them. KI-042 is narrowed (jwt maps act and azp; one issuer and no
    mapping from scopes to permissions remain).
  • Tools and the model know when another agent asks for the user. current_caller()
    returns the calling agent (Caller.actor, actor_chain, delegated); the new
    require_direct_caller() refuses unless the user asks this agent directly; require_owner
    still compares the user. require_user_mentioned follows A2A_DELEGATED_MENTIONS: origin
    (the default) also needs the id in the user's own words the calling agent forwarded, and
    refuses when none were forwarded (so, until the A2A client forwards them, a delegated
    write the check guards is refused and the user names the record at this agent directly);
    refuse always refuses; request keeps the 0.2 reading, and lint and api show point it
    out when .env or a values file sets it. In a delegated run UntrustedToolResults fences
    each human message the model reads as that agent's (<agent_request from="...">) and adds
    one factual note after the system prompt saying an agent wrote the request, with the user's
    own words when forwarded; A2A_CALLER_NOTE=off drops the note. A bad value of either
    setting stops startup. The fake test model reads the request inside that fence.
  • Approvals relayed across agents. An approval rule may let named agents deliver the
    requester's decision from another agent: decide_with: relayed with relayers (actor
    ids) in api-policy.yaml, written by api approval NAME --decide-with relayed --relayers concierge (a loosening, reviewed like new approvers; --decide-with direct narrows it
    again). The default stays direct: the person decides with their own credentials. A relayed
    decision is accepted only from a listed agent, on a thread it started for that user, with
    requester an approver, naming the approval's digest (a SHA-256 of the call as the
    approver saw it; missing or different: 409 approval_digest_mismatch), and is recorded as
    decided_via. How the approvers decide is bound to the approval when it is asked, as the
    approvers are: a policy that starts or stops relaying, or changes the relayers, while a call
    waits does not keep its approval. The approval object shows decide_with, decided_via
    and digest; the HTTP decision body and the A2A decision part take an optional digest.
    Rules that decide differently are different gates for the rule-conflict check, api show
    and lint print approved by requester; relayed by concierge, and api show --json adds
    decide_with and relayers to a relayed gate. The approvals table gains decide_with,
    relayers, decided_via and display_digest at startup, and the langgraph dev approvals
    file moves to version 2 (a version-1 file is read, its approvals direct).
  • A relayed decision shows what it decides and what will happen. An approval of an A2A
    message that approves another agent's approval carries nested (that approval, as the
    message sends it: its call, reason, expiry, digest, and the approval it relays in turn) and
    effect (the call that will actually happen, the agent that makes it and the agents via
    which), in /chat, GET /approvals, the thread's approvals and the A2A approval request. It
    expires 5 s before the approval it decides at the latest, and the nested calls' and the
    effect's query and body are dropped on decision with the call's own (unless
    TRACE_CAPTURE=full). approvals list and run print the effect first ("orders (via
    billing) will POST /orders/ORD-1002/cancel (cancelOrder), as reported by orders", its body,
    and a via line per agent), terminal-safe like the rest of the approval. With
    requester_actor and decided_via on every approval, this narrows KI-009 (approvers still
    see the requester as a hash).
  • The A2A server speaks to agents calling for a user. An agent's card declares the
    graph-agents-cli origin extension
    (https://ss7172.github.io/graph-agents-cli/a2a/ext/origin/v1, optional): an agent calling
    for a user may put the user's own words in the message metadata under that URI (origin:
    text, truncated, hops), and for a delegated caller only they reach the run's private
    credentials (@origin, where require_user_mentioned and the model's note read them),
    capped at A2A_ORIGIN_MAX_CHARS (4000). They are never stored: every task is saved without
    them. The run a decision resumes acts on the words of the request that paused it, whatever
    words the decision carries (the person's "yes, go ahead" at the agent that relays it, or
    none): the approval keeps them while it waits (never shown, and dropped once it is decided or
    expired, whatever TRACE_CAPTURE says; fastapi runtime only, as LangGraph Server passes no
    credentials to tools), so require_user_mentioned holds again on the resumed run and an
    agent relaying the approval one level further rebuilds the very call the person approved.
    More hops than AUTH_MAX_DELEGATION_DEPTH fails the task (delegation chain too deep). A
    task waiting for approval carries approval_json, the approvals as exact JSON text, and its
    text shows each call's body (up to 2,000 characters) (KI-026). A failed task, or a refused
    decision, carries a data part {"type": "error", "code": ...} (thread_busy,
    approval_direct_only, ...). A decision may be sent on the context alone, naming the waiting
    task in referenceTaskIds.
  • A2A tasks follow their approval's outcome (KI-025). However an approval ends (a decision
    over A2A, on the task or on its context; one taken over HTTP, the person at this agent
    included; or its expiry), the requester's input-required tasks on that thread that wait on
    it take the resumed run's outcome (completed, failed, or input-required with the new
    approvals) and say where it continued (Continued in task <id>., ... was approved outside this task; the run continued there., ... expired before anyone decided.), with the run's
    reply. Both task stores; another principal's task on the thread is left alone. role: gates
    are decided over HTTP, not A2A: documented as a design choice.
  • Agents call other agents for the user with a token exchanged for theirs (auth: exchange,
    RFC 8693).
    An API declared auth: exchange with exchange: {audience, scope, resource} is
    called with Bearer <token> (in forward_header, default Authorization): a token the
    identity provider mints for that audience in exchange for the caller's own verified token,
    naming this agent as the actor. It is asked for just before the call is sent
    (app_utils/token_exchange.py), after the policy check, the approval gate and the limits: a
    refused call, or one paused for a person's approval, never exchanges, and nothing is
    exchanged while a request is authenticated. Tokens are kept in process memory only, per user
    token, audience, scope and resource, for at most expires_in, 300 s
    (TOKEN_EXCHANGE_MAX_TTL_S) and the user's own token's expiry, less 30 s; concurrent calls
    share one exchange. A user token with 10 s or less left is not exchanged; an issuer refusal
    is remembered for TOKEN_EXCHANGE_FAILURE_TTL_S (10 s); three issuer failures in a row
    (timeouts after TOKEN_EXCHANGE_TIMEOUT_MS, 2 s; connection errors; 5xx; unusable answers)
    open a circuit breaker that fails calls to exchange APIs at once, then lets one call probe
    the issuer. Nothing is sent when there is no token to exchange (a shared-bearer caller, a
    run resumed by a role approver), and the tool reads why. The issuer and this agent's client
    are TOKEN_EXCHANGE_URL (https outside dev unless loopback or
    TOKEN_EXCHANGE_ALLOW_HTTP=true), TOKEN_EXCHANGE_CLIENT_ID, TOKEN_EXCHANGE_CLIENT_SECRET
    (client_secret_basic, or client_secret_post through TOKEN_EXCHANGE_CLIENT_AUTH);
    outside APP_ENV=dev the app refuses to start with an exchange API and any of the three
    missing, or under shared-bearer or langgraph-server, and a malformed setting stops it
    everywhere. Exchange APIs receive the request's X-Request-ID and trace context, as auth: forward ones do (the owner's decision). A call to an agent already in the request's
    delegation chain, or to this agent itself, is refused before anything is sent. New metrics:
    agent_token_exchanges_total{api, outcome} (issued, cached, refused, no_actor,
    unavailable, circuit_open) and agent_token_exchange_duration_seconds{api}; one log line
    per exchange sent and a warning when the breaker opens, never a token. An exchanged token
    that names no actor is refused, and nothing is sent
    (the owner's decision of 2026-09-28): a
    JWT without the act claim (or the one AUTH_JWT_ACTOR_CLAIM names; some issuers, Keycloak
    among them, add none), or a token the agent cannot read as a JWT (opaque, encrypted). The
    called agent would read such a token as the user's own, and could let this agent decide the
    user's approvals there. The refusal is remembered for TOKEN_EXCHANGE_FAILURE_TTL_S (metric
    outcome no_actor). An API opts in with exchange.allow_actorless: true, for a called agent
    that sets AUTH_JWT_DIRECT_CLIENTS and lists this agent as client:<its client id> in
    AUTH_ALLOWED_ACTORS; the first such token then logs a warning (KI-149). jwt now also
    keeps the token's exp for this (keep_subject_token takes exp). The authentication guide
    has a Keycloak recipe. Existing projects get the runtime from scaffold upgrade (a new
    app_utils/token_exchange.py; api_client.py, auth.py, metrics.py, telemetry.py,
    fast_api_app.py).
  • auth: forward can forward the caller's own token to an agent it was minted for:
    forward_audience.
    Under jwt, where no per-API credential is set, an auth: forward API
    with forward_audience sends Bearer <the caller's token> only when that token's aud
    names the audience; a token minted for this agent alone is never replayed at another. Prefer
    auth: exchange.
  • api add --auth exchange --audience AUD [--scope S] [--resource URI] [--allow-actorless]
    declares an exchange API (--allow-actorless writes exchange.allow_actorless: true);
    --audience with --auth forward writes forward_audience, and --forward-header goes
    with either mode. The first exchange API adds TOKEN_EXCHANGE_CLIENT_SECRET to the
    manifest's secrets.keys, TOKEN_EXCHANGE_URL and TOKEN_EXCHANGE_CLIENT_ID to
    .env.example (the secret commented out) and to the chart's values.yaml (a placeholder URL
    and the project's name), and lists what is left: the issuer's permission to exchange for the
    audience, and the callee's AUTH_JWT_AUDIENCE and AUTH_ALLOWED_ACTORS; without the opt-in,
    that the issuer must name this agent in act (else calls fail), and what the opt-in needs;
    with it, the callee's AUTH_JWT_DIRECT_CLIENTS and this agent as client:<id> in
    AUTH_ALLOWED_ACTORS. lint notes the same for every API that opts in. api remove takes
    them away with the last exchange API, and create --api-policy with an exchange API renders
    the same. api show names the audience, scope and resource, and the opt-in. The schema
    accepts forward_header with exchange as well as forward (both SHARED copies; the
    messages name every mode), and create, lint and api add refuse auth: exchange under
    langgraph-server as they refuse auth: forward.
  • JSON-RPC APIs and other agents in api-policy.yaml: protocol, rpc_method and
    a2a_operation.
    An API may set protocol: jsonrpc (a JSON-RPC 2.0 API) or protocol: a2a
    (another agent over A2A 1.0 JSON-RPC, with its endpoint in a2a: {path: /a2a/<name>}), and
    any API a description. For those, every POST must send one JSON-RPC request object (a
    batch, a notification, extra members, a body that is not plain JSON, or a body on a GET or
    HEAD is refused before sending), and the policy client reads what it is from the body, never
    from the tool's label: its method (rpc_method; under a2a an A2A 0.3 name such as
    tasks/cancel is read as its 1.0 name, CancelTask) and, for an A2A message whose parts
    name an approval, what it decides (a2a_operation: reject only when every such part
    rejects, else approve; a message method in any letter case is read for a decision).
    Operation entries may pin rpc_method and a2a_operation: an allow must match them; a
    denial or gate covers every call they describe whatever its path, label or spelling. A
    protocol: a2a API that can send messages must gate or deny a2a_operation: approve, so an
    agent never decides on its own an approval the agent it calls waits for (the policy is
    invalid otherwise, and the client refuses such a message at runtime too), and refuses auth: none; JSON-RPC APIs allow GET, POST and HEAD only. A tool's operation_id that names an
    entry pinning another method or decision is refused (the label cannot hide a request). Plain
    http APIs, and every entry without the new keys, are judged exactly as in 0.2 (property
    tests against the 0.2.0 rules). With protocol: a2a, the transport rule for peers applies:
    outside APP_ENV=dev a credential goes to such an API over https only, unless the host is
    loopback, a single-label or a .svc name (<ENV> must use https outside APP_ENV=dev to carry credentials); other APIs keep their transport. Both SHARED copies; every message is in the
    schema reference. KI-122 is closed for these APIs: their gates name rpc_method or
    a2a_operation, which no label can hide.
  • lint and api for JSON-RPC APIs and A2A peers. API_CALLS entries take rpc_method
    and a2a_operation: every POST to a protocol: jsonrpc|a2a API declares its rpc_method
    (an A2A 0.3 name is read as its 1.0 name), and lint judges the declaration as the client
    judges the body, refusing the keys on an http API or a GET, an operation_id that names
    another request, and a message that approves with nothing holding it; a refused JSON-RPC call
    gets its api allow ... --rpc-method M hint. lint warns about an A2A peer without a
    description and about more than 40 peers. New flags: api add --protocol http|jsonrpc|a2a --a2a-path P --description T (an a2a API that allows POST is written with
    denied_operations: [{a2a_operation: approve}], fail closed, and the note says how to gate
    the person's decision instead; --auth none is refused), api allow NAME [OPID] --rpc-method M --method POST --path P, api deny NAME --rpc-method M / --a2a-operation approve|reject,
    api revoke with the same two, and api approval NAME --a2a-operations approve[,reject]|none
    (which keeps the rule's --operations entries, and the other way round). api show prints
    the description and the protocol, and --json adds description, protocol and a2a per
    API and rpc_method and a2a_operation per declared call.
  • A decision is bound to the JSON-RPC request it was taken for. On a protocol: jsonrpc|a2a API every call is a POST to one endpoint, so API, method and path could not
    tell a relay's calls apart: the read a resumed tool sends first would have taken the
    decision its approve message waits for, and a rejection would have stopped the reject
    message that tells the other agent. The call's identity now includes its JSON-RPC method and
    A2A decision (read from the body) for such APIs, in the interrupt, the approval record (its
    payload), the decision and the ledger's bound approvals; the approval object and its
    digest include them. Calls to http APIs keep their three-field identity, and their
    records, decisions and digests are unchanged.
  • Requests keep one request id and one trace across agents (PROPAGATE_TRACE_HEADERS).
    An agent that called another agent started a new request id and a new trace there, so a
    request across agents could not be followed from end to end. Under
    PROPAGATE_TRACE_HEADERS=peers (the default), every call through the policy client to
    another agent (protocol: a2a, whatever its auth) or to an API that acts for the calling
    user (auth: forward, auth: exchange) carries the request's X-Request-ID and, under
    OTLP tracing, its W3C trace context (traceparent, tracestate). Other APIs (auth: bearer or auth: none over http or jsonrpc) are third parties and never receive them.
    An incoming traceparent is continued on the A2A routes (/a2a/*) only, so a caller of
    the public routes (/chat, the thread and approval routes) cannot choose the agent's trace
    ids; its X-Request-ID is still taken and echoed. all sends them to every API and
    continues a trace on every path (an agent behind a tracing gateway); off neither. true
    reads as peers (logged once) and false (also 0, no) as off; any other value stops
    startup. Under shared-bearer and langgraph-server, which refuse forward and
    exchange, only an agent's peers receive them. The headers are not bound by approvals.
    Existing projects get it from scaffold upgrade (app_utils/telemetry.py,
    middleware.py, api_client.py, fast_api_app.py).
  • An A2A client in the template: app_utils/a2a_client.py (B3; replaces the
    experiment's hand-written peers client). peer_tools(PEERS) gives the model
    ask_agent(agent, request) (its description lists the peers and what each does) and, for
    peers it relays approvals to, approve_agent_action(agent, task_id); A2APeerClient does
    the same from your own tools (send, get_task, pending_approvals, decide, cancel,
    relay, card). Every request goes through the API policy (the allow-list, the approve
    gate, auth: exchange, the limits, the response cap); the SDK's HTTP client is never used.
    Before the first call a peer's agent card must offer an A2A 1.x JSON-RPC interface at the
    URL this agent calls, named after a2a.path (cached A2A_CARD_TTL_S, 300 s; its URL never
    dialed). The contextId is a UUID keyed with PRINCIPAL_HASH_SALT, one per thread, peer and
    user; calls to one peer in a thread are serialized; a busy peer is asked again 3 times, and a
    cancel another replica runs once. Replies are the last response artifact (at most
    A2A_REPLY_MAX_CHARS, 6000). With A2A_FORWARD_ORIGIN=auto (default) a peer whose card
    declares the origin extension gets the user's own words. Calls back to this agent or up the
    delegation chain are refused. The relay reads what the peer waits on from the peer (its
    exact approval_json, or its approvals ledger when it lost the task), reports a gate the
    person decides at the peer as needs_direct_approval, and otherwise sends one
    context-addressed decision, the same on every run, through this agent's approve gate: the
    person approves it here, seeing the peer's call as effect; a rejection tells the peer at
    once. .env.example and the environment reference list the new settings, which the startup
    check validates. The fake test model fills a one-value Literal argument (a JSON-schema
    const).
  • graph-agents-cli peer add|remove|list|show|sync: declare the agents this agent asks
    (B3). peer add NAME writes the peer's protocol: a2a API (its endpoint /a2a/NAME, the
    credential the auth policy calls for: jwt exchange for audience NAME, custom forward,
    shared-bearer bearer; the card, SendMessage, GetTask, optionally CancelTask; the approve
    gate with --approvals relay, the default, or its denial with deny; 12 calls a run, a 120 s
    read timeout and a 1 MiB answer cap), the manifest's secrets.keys, .env.example, the
    chart's values.yaml and, with --cluster-url, each values-<env>.yaml (never .env), and
    regenerates <agent_dir>/tools/a2a_peers.py: data only (PEERS, a literal API_CALLS of
    exactly the calls the policy allows, TOOLS = peer_tools(PEERS)), importing the manifest's
    agent directory. It prints what is left, including the peer's side: AUTH_JWT_AUDIENCE and
    AUTH_ALLOWED_ACTORS there, AUTH_JWT_DIRECT_CLIENTS and client:<id> for an issuer whose
    exchanged tokens name no actor (--allow-actorless), and the api approval <its gated API> --decide-with relayed --relayers <this agent> line that lets it relay (a reviewed loosening
    the peer's owners run). Guards: a 0.2 runtime (run scaffold upgrade), a bad or own name
    (exit 2), a taken API name, a peer with other settings, an auth mode the project cannot serve,
    a module it did not write (exit 3); --card URL|FILE reads the description, path and origin
    support. peer remove takes it all back, peer list and peer show [--check] report (the
    check reads the card without a credential: reachable, 401, or a foreign endpoint, exit 1),
    and peer sync regenerates the module after scaffold upgrade or api edits. lint fails
    while the module and the policy differ, and notes a tool module that calls a peer itself.
    Wiring the round-1 concierge's five peers takes five peer add commands instead of 40 api
    commands and 418 hand-written lines.
  • deploy checks the agents the project calls before building (the peer pre-checks).
    Outside dev (the environment or the pods' APP_ENV) it refuses, exit 3, in every CD mode:
    a credential for a protocol: a2a API sent over plain http to a host that is not loopback, a
    single-label name or a .svc name (the runtime's transport rule); an auth: exchange API
    without TOKEN_EXCHANGE_URL or TOKEN_EXCHANGE_CLIENT_ID in the chart env; a plain-http
    TOKEN_EXCHANGE_URL to a host that is not loopback without TOKEN_EXCHANGE_ALLOW_HTTP. It
    warns about an unset peer URL and about TOKEN_EXCHANGE_CLIENT_SECRET or
    PRINCIPAL_HASH_SALT missing from secrets.keys; in dev the refusals warn too. A test keeps
    the CLI's copy of the runtime's internal_host equal to the template's.
  • graph-agents-cli system check|apply|graph|delegations|deploy: agent projects that call
    each other, seen as one
    (B13). An optional graph-agents-system.yaml (found upward, or
    --file; JSON Schema schemas/graph-agents-system.schema.json, generated from the models)
    names each agent's project, its client id and actor id (actor_id: the act.sub of its
    exchanged tokens, which the agents it calls list and --relayers names; default the client
    id, for an issuer that names the client otherwise there, such as agent:<client>, where
    listing the client id would leave every call refused with 403 while system check passed:
    found by the round-3 acceptance run), the agents it calls (with approvals, auth,
    scope, calls, description), the environments (port_base for local processes, a url
    template, or in-cluster URLs from each chart and manifest namespace), the token issuer, a
    shared database's max_connections and deploy.parallel. A file that cannot be used exits 3
    (an unknown project or one two agents name, an edge to an unknown agent or to itself, one
    agent called twice by another, two agents with one client id or actor id, an exchange edge without
    identity, an environment a manifest does not know). system apply writes both sides of
    every edge, idempotently and one diff per project: in each caller what peer add writes, and
    per environment the peer URLs, TOKEN_EXCHANGE_URL and networkPolicy.egressTo to the
    called agent's pods (TOKEN_EXCHANGE_CLIENT_ID is the file's client id); in each called
    agent its appUrl per environment, AUTH_JWT_AUDIENCE when empty, the callers' actor ids
    in AUTH_ALLOWED_ACTORS, and networkPolicy.ingressFrom for the callers' pods (plus the
    Gateway's namespace while the route publishes paths). It never writes gates (it prints the
    api approval ... --decide-with relayed line a relay needs), secrets, .env or local
    settings, removes a peer of the file's agents that left calls, keeps an existing peer's
    tuning, and never takes access away. system check runs SC01-SC13 (runtimes and charts,
    edges in step, paths, issuer and audience, appUrl against the URL dialled, replicas with
    in-memory tasks, relays the called agent's gates refuse, allowed actors, cycles and
    delegation depth, a shared database's connection budget, the callers' secrets, exchange under
    langgraph-server, an A2A path still public), and with --live SC14-SC15 (Services with a
    ready endpoint, read from their EndpointSlices since the v1 Endpoints API is deprecated;
    cards and the token URL answering; the Secrets' key names), using each project's recorded
    kube context and never the current one outside dev; exit 1 on an error, --json. system graph draws the system (mermaid, dot or json); system delegations prints what the issuer
    must allow each client and guarantee. system deploy --env ENV checks, then runs
    graph-agents-cli deploy in every project in waves, callees first (a cycle broken in file
    order), at most --parallel at once (default 3), stopping after a failed wave unless
    --keep-going, printing each agent's build, load and rollout times, then checks --live;
    outside dev it requires every project's recorded context, and projects in argocd mode (a
    commit and a pull request each) go one at a time. api/_files edits now build on the planned
    text of a file, so one plan can hold several peers. The NetworkPolicy rules select an agent's
    pods by the chart's selector labels (name and release: the bundled database shares the
    release label), and were checked on a kind cluster with kindnet enforcing them: callers reach
    a called agent's Service port 80 through an egress rule on the pods' port 8000 (a rule on
    port 80 blocks them), and a pod in another namespace is refused. KI-159 and KI-160 park the
    residuals (LangGraph Server's own pool in SC10, a local environment's .env unchecked);
    KI-158 (SC14 on the deprecated Endpoints API) was fixed before release.
  • A guide to agents calling agents
    (website/src/guides/multi-agent.md): who acts for whom,
    a walk-through from peer add to an eval at the entry agent, relayed approvals, the user's
    own words, the system view, following one request across agents, sizing, the threat model
    with what remains, and the limits. The security checklist adds the internal A2A paths, what
    the issuer must allow and the salt for agents that ask others; the observability guide names
    the actor log field; the manifest reference says who adds TOKEN_EXCHANGE_CLIENT_SECRET
    and peers' keys to secrets.keys. A reference page,
    graph-agents-system.yaml, lists every key of the
    system file with its default and what makes a file unusable.
  • The skills cover agents calling agents. The workflow skill asks in Phase 0 whether the
    agent asks other agents or is called by them, declares peers with peer add or system apply (never a hand-written client), runs lint, peer show --check or system check,
    evaluates at the entry agent and deploys with system deploy. The langgraph-code skill has
    a section on the generated tools/a2a_peers.py and A2APeerClient, what tools see when an
    agent asks for the user (current_caller, require_direct_caller, require_owner,
    require_user_mentioned under A2A_DELEGATED_MENTIONS) and relays (decide_with). The
    deploy skill covers system apply|check|deploy|delegations, AUTH_ALLOWED_ACTORS, appUrl
    per environment, keeping A2A paths internal and sizing a shared database; the scaffold skill
    says peer add writes the client side and that langgraph-server calls other agents only
    with auth: bearer. Each changed skill passes gac-bench's fact-check (every command and
    option exists; within 1.25 times the 0.2.0 text).
  • limits.max_response_bytes caps an API's answers. With it (1 to 67108864 bytes), the
    client reads a response body, decoded, only up to that many bytes: past it, the answer is
    discarded and the call fails with <api> answered with more than N bytes; discarded (a
    declared Content-Length over the cap is refused before reading). Memory stays bounded
    whatever the answer's compression: a capped call asks for gzip or deflate at most and
    decodes the body itself, never past the cap (httpx decodes each network read whole, and 32
    KiB of zstd is 1 GiB), and an answer in any other content encoding (zstd, br, an unknown or
    a stacked one) is refused unread: <api> answered in a content encoding other than gzip or deflate, which limits.max_response_bytes cannot bound; discarded. api add --max-response-bytes N and api limits NAME --max-response-bytes N|none set it, and api limits edits keep it; api show prints it. Unset, answers are read whole as in 0.2 (no
    default cap for existing APIs). Both SHARED copies.
  • Reasoning effort and the Responses API for OpenAI-API models: MODEL_REASONING_EFFORT
    (none, minimal, low, medium, high, xhigh) and MODEL_USE_RESPONSES_API (true:
    every request to /v1/responses; false: Chat Completions; unset: langchain-openai
    chooses, as before), for openai and openai-compatible; the judge reads
    JUDGE_REASONING_EFFORT and JUDGE_USE_RESPONSES_API, defaulting to the agent's when it is
    an OpenAI-API model too. A model that refuses function tools with a reasoning effort on Chat
    Completions (the experiments' finding F15: its first call failed with a 400 naming
    /v1/responses) now works with the switch on, tools, streaming and token usage included. A
    bad value, or either setting for another provider, stops startup. .env.example and, for
    OpenAI-API projects, the chart's values.yaml document them; existing projects get the
    runtime from scaffold upgrade (app_utils/model.py, fast_api_app.py).
  • Structured final answers: a JSON schema for the agent's answer. A project that puts a
    JSON Schema (root "type": "object") in <agent directory>/response_schema.json
    (create --response-schema FILE seeds it) has its agent answer in that shape. agent.py
    builds the agent with response_format=response_format(model, tools)
    (app_utils/structured.py), one of LangChain's create_agent strategies:
    RESPONSE_FORMAT_STRATEGY=auto (default) uses the provider's own structured output where
    LangChain's model profile says the model has it with the agent's tools bound and the
    model's client can send the schema, strict on OpenAI (LangChain's own auto mode asks OpenAI
    for a best-effort schema), else a final_answer tool the model must call (tool_choice
    forces a call at every step); provider and tool force one. Anthropic's client refuses a
    type list and a schema with no type (an enum alone) before any request, so auto uses
    the tool for such a schema and provider stops startup naming why. LangChain returns a raw
    JSON-schema answer unchecked, so the new StructuredAnswer middleware (last in
    middleware()) checks every answer against the schema: one that does not fit, a reply that
    is not JSON, a plain-text final reply, or an answer given beside other tool calls (none of
    which runs, so a gated call never runs after an answer already given) goes back to the model
    with what is wrong, up to 3 tries in the same step (the failed tries stay out of the thread;
    their tokens count in the answer's usage), then the run ends with the new error code
    invalid_structured_response. The checker supports a documented JSON
    Schema subset and refuses a schema that uses anything else at startup (lint and
    create --response-schema apply the same rules, a SHARED block kept byte-identical with the
    template's). Delivery: a completed /chat run's only message.delta is the answer's JSON
    text and message.end carries the object as structured_response; the answer tool never
    shows as a tool.call; a run paused for an approval answers once resumed; the A2A
    response artifact adds a data part with the object (mediaType application/json,
    streamed or not; an A2A task that waited on an approval decided elsewhere, over HTTP
    say, takes the whole answer as its last response artifact) and the card lists
    application/json among its output modes; eval
    records structured_response in its traces and expect.json_schema checks it as it is.
    Both runtimes, both checkpointers. Without the file nothing changes. Measured with
    gpt-5-mini on a 24-case triage task (a tool call, then a six-field answer), twice: without
    the mode 28 of 48 replies were not a bare JSON document (the sentence the model writes
    before a tool call came first), with it 0 of 96 (both strategies, no correction needed),
    at the same cost for the provider strategy and about twice the output tokens for the tool
    strategy. Existing projects:
    scaffold upgrade brings structured.py and the runtime, but never rewrites agent.py:
    pass response_format=response_format(model, tools) to create_agent and add
    StructuredAnswer() last to middleware() before adding a schema (lint warns about
    either). The runtime checks every answer again before delivering it, so an answer that
    never went through StructuredAnswer and does not fit ends the run with
    invalid_structured_response and is not sent on /chat or over A2A (it stays in the
    thread, KI-172). Project tests: a generated project's tests run with the mode off
    (RESPONSE_SCHEMA_PATH=none, a new value: no schema, whatever the file), so a project with
    a schema keeps a green suite and CI, and tests/unit/test_structured.py checks the
    project's own schema and that one turn of agent.py answers in it; the fake model takes a
    choice of the schema that fits (null for the develop guide's optional order_id) where
    its text breaks a pattern. Documented in
    Develop your agent, the HTTP API
    and environment references, the upgrading guide (0.2 to 0.3) and the multi-agent guide (a
    peer that answers in JSON); the langgraph-code, workflow, scaffold and eval skills teach it
    (declare the shape, never parse JSON out of a reply, wire a 0.2 agent.py by hand); and
    gac-bench has a task family for it. KI-165 to KI-177 park its remaining minor issues
    (KI-168 also covers A2A).
  • A documentation site in website/ (MkDocs Material): Get started (installation, a
    five-minute quickstart, two tutorials, the lifecycle), guides for building and operating an
    agent, and a reference whose CLI and Skills pages are generated from the commands and
    skills/. .github/workflows/docs.yml builds it with mkdocs build --strict and checks its
    links on every pull request, and publishes it to GitHub Pages from main when the
    repository variable PUBLISH_DOCS is true: the site is at
    https://ss7172.github.io/graph-agents-cli/. Preview it with
    uv run --group docs mkdocs serve -f website/mkdocs.yml.
  • gac-bench, contributor tooling for the skills (tools/skillopt/): a benchmark of
    realistic graph-agents-cli tasks for each of the six skills, 104 of them, including this
    release's agent-to-agent features (peer add, relayed approval gates, auth: exchange,
    rpc_method rules and system apply) and structured final answers, each with a
    deterministic verifier and scripted gold and broken solutions, in frozen train, val and
    test splits; and
    gac_skillopt, an environment in which SkillOpt
    runs a candidate skill in Claude Code or Codex, isolated so that the session sees only that
    skill, scores it and proposes edits. It found the skill rules this release adopts after
    review, and tools/skillopt/results/ keeps every measurement. It is never a dependency of
    the CLI or of a generated project and is in neither the wheel nor the sdist (a fast test
    guards the build configuration); CI runs its unit tests. CONTRIBUTING.md and the site's
    Skills benchmark page say how to run it.

Changed

  • The README is a short entry point with absolute links (it is also the PyPI page). Its
    former sections, including "Known limitations" and "Where it is behind", moved to the
    site's pages: each limitation now sits on the page of the feature it concerns, the
    comparison on Compared with google-agents-cli. The
    tests that ran README examples now run the same examples from the site's guides.
  • The workflow and scaffold skills carry rules found by the SkillOpt experiment.
    SkillOpt optimised the skills against gac-bench, a
    benchmark of graph-agents-cli tasks carried out by Claude Code and Codex sessions. Its
    proposals were reviewed by hand (wording taken from the benchmark generalised, one claim
    corrected, a misplaced rule moved, a rule learned from the benchmark's harness dropped) and
    measured again before they were adopted. The workflow skill now scopes Phase 0 and its spec
    gate to a new agent: a concrete change to an existing project (add a retry to a tool, bump
    a dependency, fix a crash) is not a new agent, and is made without a new spec, while the
    decisions it raises that the user owns (new API operations or access, approval gates, the
    model) still go to the user. This rule was written by hand after an earlier candidate made
    Codex refuse such changes. For a new agent it says what counts as approval of the spec: a
    request to build, a draft spec, defaults the agent chose or an instruction to proceed on its
    own is not one. With nobody to approve, the agent stops before create at a draft spec and
    ends its answer with each open decision as a direct question, and asks for approval as a
    question. It also says to fix the agent, not the eval, when an eval that passed breaks
    (unless the user asked for the change the eval checks), maps failed checks to code, says how
    to show that a failing test is unrelated to a change, proves a change with lint, the eval run counts and a run smoke test, and never presents the fake model's exit 0 as evidence
    of quality. On the benchmark's workflow val tasks, against the 0.2.0 text: with Codex
    (gpt-5.6-terra, 2 repetitions) 12 of 12 passed against 6 of 12, the open decisions were
    asked as questions in 6 of 6 spec-gate sessions against 0 of 6, and every requested change
    was still made (6 of 6); with Claude Code (3 repetitions) 17 of 18 passed against 10 of 18,
    the spec-gate sessions stopped before create in 9 of 9 against 1 of 9 and asked in 9 of 9
    against 1 of 9. The adopted wording, with generic examples, was checked again on Claude
    Code: 18 of 18, with changes made, stops and questions each 9 of 9. The scaffold skill keeps
    the default AGENTS.md guidance file unless the user says the team uses one coding agent
    (8 of 8 sessions, against 4 of 8), its examples no longer pass
    --agent-guidance-filename CLAUDE.md, and it has the agent read create_params before
    scaffold enhance and after create.
  • The observability skill has a procedure for salting the hashed principal id, found by
    the SkillOpt experiment with Codex sessions and reviewed by hand before it was adopted: add
    PRINCIPAL_HASH_SALT to secrets.keys (configuration only: no code, chart template or
    values change), put the value in .env.<env>, then secrets apply --env <env> and
    deploy --restart --env <env> for every deployed environment, and name those commands
    verbatim when the user runs them. Codex followed it in 6 of 6 sessions of the benchmark's
    salt task, against 1 of 9 with the old text. Tracing for one environment now edits only
    that environment's files and ends with the commands still to run; local tracing edits
    only .env and adds none of the LangSmith SDK's own switches. Its paragraph on trace
    headers across agents states this release's rule (PROPAGATE_TRACE_HEADERS=peers|all|off:
    other agents and auth: forward/exchange APIs, and an incoming trace continued on
    /a2a/* only) instead of 0.2's.
  • Approval times read from Postgres are in UTC, whatever the database session's time zone
    (created_at, expires_at, decided_at, used_at; they came back in that zone, the same
    instant written differently). An approval now reads back with the very times it was
    created with, which the A2A client's relay binds in its decision when it reads what a peer
    waits on from the peer's approvals ledger (the peer lost the task): on a database not set
    to UTC that decision differed from the one the person approved, and nothing was sent.

Fixed

  • eval's json_schema check reads the reply's answer, not its first JSON. It parsed the
    first fenced code block, and otherwise everything from the first { or [ to the end of
    the reply: a reply that showed an example (or quoted its input) before its answer was
    checked against the example, and one that added prose after raw JSON failed as "not valid
    JSON" (the experiments' finding F10). It now reads the whole reply when that is JSON, else
    the reply's last JSON object or array, taking the schema's root type when it names
    object or array (so a citation such as [1] after an object answer is skipped), inside
    code fences or not.
  • run, eval run, eval generate and approvals no longer fail under a SOCKS proxy.
    With ALL_PROXY=socks5h://... in the environment (as coding-agent sandboxes such as Codex's
    network proxy set it), every request crashed with ImportError: Using SOCKS proxy, but the 'socksio' package is not installed, even to the command's own local server and even when
    NO_PROXY listed it: httpx builds a transport for every proxy variable when a client is
    created. Requests to this machine (the local server, the playground) now never use the
    environment's proxies. Requests to another machine still honour HTTP_PROXY, HTTPS_PROXY
    and NO_PROXY, and SOCKS proxies now work too: graph-agents-cli depends on httpx[socks]
    (adds the pure-Python socksio). A proxy setting httpx cannot use (another scheme) is a
    one-line error naming the variable, exit 3, instead of a traceback; eval generate --url
    reports it once, before any case runs.
  • The local server is stopped even where listing processes is denied. In a sandbox that
    refuses the process table (Codex's seatbelt denies kern.proc.all), psutil raised
    PermissionError while run, eval run or eval generate stopped the server they
    started: the command ended with Error: PermissionError: [Errno 1] Operation not permitted and the server kept its port, orphaned. The teardown now signals the server's
    own process group (it is started in a new session, and its uvicorn child stays in it) and,
    for a server this invocation started, its process handle; no error from the teardown
    replaces the error of a failed start any more.
  • run --stop-server no longer reports success for a server it could not stop. When the
    operating system refused the signal (Claude Code's sandbox lets a command signal only the
    processes it started itself, so a server started by an earlier command survives), it
    printed "Local server stopped." and deleted the record while the server kept its port; the
    next eval run then failed with "something is already listening". It now exits 2, names
    the processes still running and how to stop them, and keeps the record, so later runs reuse
    that server. A server that run, approvals or eval generate cannot stop after their
    work is a warning that never replaces their own result or error. Neither does a server a
    command would replace (idle for 30 minutes, or no longer answering) but may not stop: a
    warning names it and the kill command, then the command reuses it while it still answers,
    or starts a fresh one beside it.
  • info honours GRAPH_AGENTS_CLI_NO_UPDATE_CHECK=1 and CI. It ran npx -y skills@1.5.9 list --json every time, which may download the package: a disconnected
    install waited up to 15 s for it, and in a coding agent's sandbox the blocked download read
    as "the CLI is blocked". With the variable set (the disconnected profile) or a CI marker, the
    listing is skipped like the skills version check, and info says so ("Installed skills:
    not listed (... skips npx skills list)"; --json adds installed_skills_skipped with the
    reason, null when the listing ran).
  • eval grade works in a project with another agent directory. The judge runner
    imported app.app_utils.model whatever the project's package was, so in a project created
    with create --agent-directory <name> every judge metric (the default dataset's
    response_quality included) failed with exit 3, "the judge model is unreachable or
    misconfigured". The runner now imports the judge from the manifest's agent_directory.
  • deploy works for a project that needs no Secret key. An env file that sets none of
    the allow-listed keys (a keyless project: the fake model, a keyless openai-compatible
    endpoint, no shared-bearer key) made deploy exit 3, "No allow-listed secret values to
    apply", and only after it had built the image and loaded or pushed it. Such an env file now
    leaves the Secret as it is, like a missing env file in dev, and says so; the decision is
    made before anything is built. Outside dev, where the chart requires the Secret
    (secretOptional: false), a missing Secret stops the deploy (exit 1) before the build.
    secrets apply still exits 3 when there is nothing to apply.
  • A2A tasks are shared by every replica and survive restarts (KI-024, now Low). The task
    store was in process memory, per replica, while the production values run two replicas:
    GetTask, ListTasks, CancelTask and a message naming a taskId (an approval decision)
    failed with -32001 "Task not found" whenever a request reached the other pod (7 of 30
    GetTask calls over new connections in the A2A experiment), and a restart or rollout
    dropped every task, including those waiting for an approval. Under CHECKPOINTER=postgres
    (and a Postgres DATABASE_URI under langgraph-server) tasks are now kept in the app's
    database, table a2a_tasks (agent_a2a_tasks), with the same per-principal ownership,
    A2A_TASK_TTL_S expiry (on the database clock; 0 now keeps a task until its thread is
    deleted) and deletion with their thread, on every replica. A task whose run ended with its
    process turns failed instead of staying working. A SubscribeToTask or CancelTask
    that reaches a replica other than the one running the task is refused (-32004, -32002)
    instead of waiting for events that happen elsewhere, or reporting a cancel the run then
    overwrites. Streamed reply chunks are written at most once a second per task. Postgres
    cannot store the character U+0000, so a stored task holds U+FFFD in its place (a message,
    a tool's output or a contextId holding one never fails the task's save).
    CHECKPOINTER=memory keeps the in-memory store. The a2a SDK's DatabaseTaskStore was
    not used: its Postgres extra is SQLAlchemy on asyncpg (the template uses psycopg only),
    its context_id column holds 36 characters where the template accepts 128, and it
    creates its table on first use, outside the schema lock that replicas share. Existing
    projects get it from scaffold upgrade (it changes app_utils/a2a.py and
    app_utils/db.py); the table is created at the next start.
  • A tool's output that is not valid Unicode no longer breaks the run. A lone surrogate
    in a tool's result (an upstream JSON "\ud800" escape decodes to one, and UTF-8 cannot
    encode it) ended the /chat stream at its tool.result event, and failed the A2A task
    with -32603 and the Python exception text ('utf-8' codec can't encode character ...: surrogates not allowed); under langgraph-server the run failed (run_failed). With a
    real model provider the model's next request could not be encoded either, and under
    CHECKPOINTER=memory every later turn of that thread failed the same way. Each time, the
    tool had already acted. The agent middleware UntrustedToolResults now replaces each lone
    surrogate with U+FFFD as the result leaves the tool, and the model's request, the /chat
    events, the thread history and A2A replies are sent as valid text whatever their source.
    Existing projects get it from scaffold upgrade (it changes app_utils/content.py,
    chat.py and a2a.py); an agent.py that builds its own middleware list keeps
    UntrustedToolResults in it.
  • Streamed runs on an OpenAI-compatible endpoint record their token usage. The generated
    agent now asks OpenAI-API models for usage on streamed responses
    (stream_options.include_usage); langchain-openai did so by itself only for
    api.openai.com, so with MODEL_PROVIDER=openai-compatible (a proxy, a gateway or a
    self-hosted server) run records and eval traces showed zero tokens and
    expect.max_tokens passed whatever the run used. On the Responses API
    (MODEL_USE_RESPONSES_API) usage comes with every streamed answer anyway, and nothing
    more is sent. Existing projects get it from scaffold upgrade (app_utils/model.py).
  • KI-146: the observability guide says that under shared-bearer, as under
    langgraph-server, only an agent's peers (protocol: a2a) receive the request id and
    trace context.
  • KI-111: the policy lifecycle no longer starts from a read-only example; the
    Outbound API policy guide makes the access level an
    explicit choice at every step.
  • KI-113: the documentation site is published at
    ss7172.github.io/graph-agents-cli, so the
    README's links to it work.

Security

  • Under langgraph-server, a native run no longer chooses who its tools act for. The
    server's native run API (POST /threads/{thread_id}/runs, /runs/wait, /runs/stream
    and the thread-less /runs routes; the default /threads path prefix publishes the first
    three) takes the run context from the request (context, or config.configurable, which
    the server copies into it), and tools read the calling principal from that context
    (current_caller(), require_owner, require_direct_caller, require_user_mentioned).
    The auth handler checked whose thread it was but passed the context through, as in 0.2.0:
    any authenticated user could have the tools act as another principal, with roles of their
    choosing ({"principal_id": "bob", "roles": ["ops"]}), and an agent presenting a user's
    request could drop its @actor and pass require_direct_caller as the user. The handler
    now puts the caller's own run context on every run it authorizes, whatever the request
    sent: the caller's id, its roles (an agent's are only those AUTH_DELEGATED_ROLES lends)
    and its public attributes with @actor (never credentials), the context /chat sends
    (authenticate publishes it in the server's user as run_context). A native run that
    sends none now acts for its caller, where every tool that needs one refused (KI-020 no
    longer lists that). Studio under langgraph dev keeps the context it sends. Existing
    projects get it from scaffold upgrade (it changes app_utils/auth.py and chat.py).
  • Agents that call each other for a user fail closed. A request another agent presents
    for a user is refused until the called agent lists that agent (AUTH_ALLOWED_ACTORS,
    empty by default), keeps none of the user's roles (AUTH_DELEGATED_ROLES, empty), reaches
    only the threads, tasks and approvals it started for that user, and never decides an
    approval unless the gate relays through it by name (decide_with: relayed with
    relayers, naming the approval's digest; the default is direct). A calling agent
    exchanges tokens only after the policy check and the approval gate, refuses an exchanged
    token that names no actor unless the API opts in, sends credentials to other agents over
    TLS outside dev (loopback, single-label and .svc hosts excepted), and refuses to call an
    agent already in the request's chain. A message that
    approves another agent's approval must be gated or denied by the caller's policy. These
    close the design flaws the A2A multi-agent experiment found in a hand-built system (an agent
    holding a user's delegated token could decide that user's approvals; any agent in a chain
    could read or cancel another's tasks); 0.2.0 itself had no delegated principals.