Skip to content

Releases: ss7172/graph-agents-cli

graph-agents-cli 0.3.1

Choose a tag to compare

@github-actions github-actions released this 29 Sep 23:14

0.3.1 is the first release published on PyPI.
Its commands and templates are 0.3.0's (only the version, the documentation, the skills'
install lines and the release workflow change), so a project created by 0.3.0 needs no
migration.

Changed

  • Published on PyPI. Install with uv tool install graph-agents-cli (or pipx install graph-agents-cli); the release tag git+https://github.com/ss7172/graph-agents-cli@v0.3.1
    installs the same release. 0.3.0 and earlier stay on their git tags only.
  • The README and the documentation point to PyPI: the README's install section (with a
    PyPI badge), the site's Installation page, announcement bar, home page and tutorials, and
    the workflow and scaffold skills' install lines. "Not on PyPI yet" is gone. setup,
    update, the scaffold upgrade baseline and generated projects' GRAPH_AGENTS_CLI_SPEC
    still install from the release tag, since earlier releases exist only as tags;
    GRAPH_AGENTS_CLI_INSTALL_SPEC='graph-agents-cli=={version}' switches them to PyPI.

Fixed

  • The release workflow's GitHub Release job is idempotent. A second run for the same
    tag (GitHub started two for the v0.3.0 push, and the second failed with "a release with
    the same tag name already exists") now finds the release, uploads only the assets it lacks
    and never replaces one; the PyPI job skips files already on PyPI.

graph-agents-cli 0.3.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 16:53

0.3.0 lets agents call other agents for the user they serve, over A2A.
Each agent knows the user and the agent in between, a person's approvals stay with that
person, peer add declares the agents one asks, and graph-agents-cli system checks, wires
and deploys several projects as one. It also keeps A2A tasks in Postgres so that replicas
share them, adds structured final answers (an agent that answers in a JSON shape the
project declares), reasoning effort and the Responses API for OpenAI-API models, a
documentation site, skill rules found with the SkillOpt experiment and the benchmark that
measured them, and fixes. What an upgrade from 0.2.0 changes, and the order to do it in, is
in Upgrading projects; the guide to
the new features is Agents calling agents. Parked
medium- and low-priority issues are listed in KNOWN_ISSUES.md.

Breaking changes and migration

Each change below can need an edit in an existing project; the upgrading guide's
0.2 to 0.3 section lists every other
change in behaviour.

  • jwt reads the RFC 8693 act claim. A token carrying it is an agent's for the user,
    refused (403) until AUTH_ALLOWED_ACTORS lists the agent; set AUTH_JWT_ACTOR_CLAIM=
    (empty) to read every token as the user's own, as 0.2 did. A custom policy that returns an
    invalid id (empty, over 256 characters, or with control characters) now fails the request
    with 500 and logs the bug. A custom policy through which other agents forward users'
    credentials must set Principal.actor (policies/ is not upgraded): see the upgrading
    guide.
  • An API whose auth the project's auth policy can never serve stops the app outside
    dev.
    lint and api add check each API's auth against the project's auth policy:
    auth: exchange or auth: forward under shared-bearer, and auth: forward under jwt
    without forward_audience, are errors (exit 3): such an API never had a credential to send,
    so every call to it failed with "the caller has no credential". Outside APP_ENV=dev the
    running app now refuses to start with such an API too, as it does for auth: exchange (the
    owner's decision of 2026-09-28; also auth: forward under the langgraph-server runtime);
    under dev it logs why and starts. Migration: a 0.2 project with such a forward API
    stops starting outside dev after the upgrade until the API gets forward_audience, moves to
    auth: exchange or is removed (scaffold upgrade never rewrites api-policy.yaml, and
    lint names the API). lint and api add also note a forward API with
    forward_audience under jwt ("prefer auth: exchange").

Added

  • Agents calling agents: the caller's identity. A request another agent presents for a
    user is now told apart from the user's own. Principal.id stays the user (the subject);
    the new Principal.actor names the agent presenting the request (its id, the chain of
    agents before it, and its client). jwt reads the RFC 8693 act claim
    (AUTH_JWT_ACTOR_CLAIM, nested act for earlier agents), the client from azp/client_id
    (AUTH_JWT_CLIENT_CLAIM), and with AUTH_JWT_DIRECT_CLIENTS treats a token with no act
    from any other client as that client's; a custom policy sets actor itself
    (actor_from_claims and keep_subject_token are exported for it). Every policy then goes
    through one rule set (finalize_principal): valid ids, at most AUTH_MAX_DELEGATION_DEPTH
    agents (default 3; else 401), only the agents AUTH_ALLOWED_ACTORS lists (default none;
    else 403), and only the roles AUTH_DELEGATED_ROLES lends. Threads, A2A tasks and approvals
    are owned by the subject and the actor: an agent reaches only what it started for that
    user, while the user owns everything done for them: with their own token they read,
    continue and delete the threads their agents started, decide their approvals, and read,
    list and cancel the A2A tasks those agents started for them (the owner's decision of
    2026-09-28; continuing or subscribing to such a task stays with its agent). A delegated
    principal's roles never read across, administer or decide as a role: approver, and it
    never decides an approval (403 approval_direct_only: the person decides with their own
    credentials). The actor reaches
    tools in attributes["@actor"], is recorded with each approval (requester_actor), and is
    logged (actor), kept in run records and named in trace metadata. auth dev-token mints
    such tokens locally (--act, repeatable, and --azp). The database gains
    threads.actor, runs.actor and approvals.requester_actor at startup (ADD COLUMN IF NOT EXISTS); existing rows are direct, and every 0.2 principal is direct, so 0.2 behaviour is
    unchanged for them. KI-042 is narrowed (jwt maps act and azp; one issuer and no
    mapping from scopes to permissions remain).
  • Tools and the model know when another agent asks for the user. current_caller()
    returns the calling agent (Caller.actor, actor_chain, delegated); the new
    require_direct_caller() refuses unless the user asks this agent directly; require_owner
    still compares the user. require_user_mentioned follows A2A_DELEGATED_MENTIONS: origin
    (the default) also needs the id in the user's own words the calling agent forwarded, and
    refuses when none were forwarded (so, until the A2A client forwards them, a delegated
    write the check guards is refused and the user names the record at this agent directly);
    refuse always refuses; request keeps the 0.2 reading, and lint and api show point it
    out when .env or a values file sets it. In a delegated run UntrustedToolResults fences
    each human message the model reads as that agent's (<agent_request from="...">) and adds
    one factual note after the system prompt saying an agent wrote the request, with the user's
    own words when forwarded; A2A_CALLER_NOTE=off drops the note. A bad value of either
    setting stops startup. The fake test model reads the request inside that fence.
  • Approvals relayed across agents. An approval rule may let named agents deliver the
    requester's decision from another agent: decide_with: relayed with relayers (actor
    ids) in api-policy.yaml, written by api approval NAME --decide-with relayed --relayers concierge (a loosening, reviewed like new approvers; --decide-with direct narrows it
    again). The default stays direct: the person decides with their own credentials. A relayed
    decision is accepted only from a listed agent, on a thread it started for that user, with
    requester an approver, naming the approval's digest (a SHA-256 of the call as the
    approver saw it; missing or different: 409 approval_digest_mismatch), and is recorded as
    decided_via. How the approvers decide is bound to the approval when it is asked, as the
    approvers are: a policy that starts or stops relaying, or changes the relayers, while a call
    waits does not keep its approval. The approval object shows decide_with, decided_via
    and digest; the HTTP decision body and the A2A decision part take an optional digest.
    Rules that decide differently are different gates for the rule-conflict check, api show
    and lint print approved by requester; relayed by concierge, and api show --json adds
    decide_with and relayers to a relayed gate. The approvals table gains decide_with,
    relayers, decided_via and display_digest at startup, and the langgraph dev approvals
    file moves to version 2 (a version-1 file is read, its approvals direct).
  • A relayed decision shows what it decides and what will happen. An approval of an A2A
    message that approves another agent's approval carries nested (that approval, as the
    message sends it: its call, reason, expiry, digest, and the approval it relays in turn) and
    effect (the call that will actually happen, the agent that makes it and the agents via
    which), in /chat, GET /approvals, the thread's approvals and the A2A approval request. It
    expires 5 s before the approval it decides at the latest, and the nested calls' and the
    effect's query and body are dropped on decision with the call's own (unless
    TRACE_CAPTURE=full). approvals list and run print the effect first ("orders (via
    billing) will POST /orders/ORD-1002/cancel (cancelOrder), as reported by orders", its body,
    and a via line per agent), terminal-safe like the rest of the approval. With
    requester_actor and decided_via on every approval, this narrows KI-009 (approvers still
    see the requester as a hash).
  • The A2A server speaks to agents calling for a user. An agent's card declares the
    graph-agents-cli origin extension
    (https://ss7172.github.io/graph-agents-cli/a2a/ext/origin/v1, optional): an agent calling
    for a user may put the user's own words in the message metadata under that URI (origin:
    text, truncated, hops), and for a delegated caller only they reach the run's private
    credentials (@origin, where require_user_mentioned and the model's note read them),
    capped at A2A_ORIGIN_MAX_CHARS (4000). They are never stored: every task is saved without
    them. The run a decision resumes acts on the words of the request that paused it, whatever
    words the decision carries (the person's "yes, go ahead" at the agent that relays it, or
    none): the approval keeps them while it waits (never shown, and dropped once it is decided or
    expired, whatever TRACE_CAPTURE says; fastapi runtime only, as LangGraph Server passes no
    credentials to tools), so require_user_mentioned holds again on the resumed run and an
    agent relaying the approval one level further rebuilds the very call the person approved.
    More hops than AUTH_MAX_DELEGATION_DEPTH fails the task (delegation chain too deep). A
    ...
Read more

graph-agents-cli 0.2.0

Choose a tag to compare

@github-actions github-actions released this 25 Sep 01:14

graph-agents-cli is now a generic CLI for building, evaluating and deploying LangGraph agents
on self-hosted Kubernetes, for any project and any domain. Nothing in the CLI, the template,
the skills or a generated project is shaped around one consumer: projects choose an auth
policy and declare the outbound APIs their tools may call. This release also closes most of
the production-readiness findings of an independent assessment of 0.1.0 (runtime guardrails,
per-user authentication, deploy safety, supply chain, release engineering); the remaining
ones are listed under "Known limitations" and "Where it is behind" in the README.

Parked medium- and low-priority issues are listed in KNOWN_ISSUES.md.

Install from the release tag (the package is not on PyPI yet):

uv tool install git+https://github.com/ss7172/graph-agents-cli@v0.2.0

Breaking changes and migration

  • --auth-policy product-session is now custom. create --auth-policy product-session
    is refused with a hint. A manifest or AUTH_POLICY that still says product-session is
    read as custom, with a one-line deprecation warning. In the template,
    app/policies/product_session.py (ProductSessionPolicy) became app/policies/custom.py
    (CustomPolicy). Set AUTH_POLICY=custom and auth_policy: custom.
  • The product API policy is now a multi-API outbound policy. product-policy.yaml
    becomes api-policy.yaml, create --product-policy becomes create --api-policy, the
    manifest block product_api: becomes api_policy: {policy_file: api-policy.yaml}, the
    cookiecutter variable has_product_policy becomes has_api_policy, and the tool
    declaration PRODUCT_CALLS becomes API_CALLS with an "api" key per entry. The file
    now declares any number of APIs under apis: {<name>: ...}; allowed_methods is
    required; auth: forwarded-session is now auth: forward. create, scaffold enhance,
    scaffold upgrade and lint stop on a project that still uses the old format and print
    the migration steps (exit 3); --product-policy is refused with a rename hint; a tool
    module that declares PRODUCT_CALLS is a lint error.
  • Outbound API calls fail closed. Without api-policy.yaml, or for an API the file does
    not declare, get_client() raises ApiPolicyError and nothing is sent. 0.1.0 sent every
    call unrestricted (with a warning) when no policy file existed.
  • The default guidance file is AGENTS.md (was GEMINI.md). Existing projects keep the
    file name their manifest records; pass --agent-guidance-filename GEMINI.md to create to
    keep the old default.
  • CLI_VERSION_PIN is now GRAPH_AGENTS_CLI_SPEC in .github/agent.env (the
    cookiecutter variable cli_version_pin is now cli_install_spec). The value is a full
    install spec, git+https://github.com/ss7172/graph-agents-cli@v0.2.0 by default, and the
    workflows run uvx --from "$GRAPH_AGENTS_CLI_SPEC" graph-agents-cli .... The workflows
    refuse an agent.env that still sets CLI_VERSION_PIN, with a rename hint.
  • Installation moved to a pinned git reference. The package name graph-agents-cli was
    never published on PyPI, so uv tool install graph-agents-cli never worked; setup,
    update, the scaffold upgrade baseline and the generated workflows install
    git+https://github.com/ss7172/graph-agents-cli@v<version> instead, and the update check
    reads GitHub releases. GRAPH_AGENTS_CLI_INSTALL_SPEC overrides the source (a mirror, a
    wheel); write {version} where the version goes.
  • deploy and secrets apply outside dev need an explicit env file and kube context.
    They read .env.<env> (or --env-file) and never fall back to .env (exit 3 without one).
    The kube context must be recorded as environments.<env>.context or passed with
    --context; the kubeconfig's current context is used only after a confirmation prompt, or
    --yes when there is no terminal (exit 1 otherwise). CI jobs that deploy need --yes or
    --context (the generated workflows pass both).
  • A live API_KEY is never replaced implicitly. Under shared-bearer, secrets apply
    and deploy keep the key in the cluster unless the env file sets a different one and
    --rotate-api-key is passed. A generated key is written to the env file (mode 0600)
    instead of being printed.
  • helm-push reads DEPLOY_KUBECONFIG from the staging and production GitHub
    environments
    (was the repository secret KUBECONFIG), so the production reviewers gate
    it. Create the environment secrets and delete the repository secret. deploy --env staging|prod is refused outside CI in helm-push mode even with --image
    (--force-direct overrides).
  • Chart defaults are stricter. image.tag defaults to "" in every environment and the
    chart refuses to render without a tag (never latest by default); the tag must be a
    quoted string. The HTTPRoute and Ingress publish only route.publicPaths (/chat,
    /threads, /a2a/<agent>, plus route.devPaths under APP_ENV=dev) instead of every
    path; /health, /ready and /metrics stay inside the cluster. Outside dev the app
    Secret is required (secretOptional: false): pods do not start without it.
  • Exit codes are consistent (0 ok, 1 refused or failed gate, 2 tool failure, 3
    configuration error): an unexpected crash exits 2 (was 1), running outside a project
    exits 3 (was 1), run exits 2 when the agent cannot be reached or goes silent (was 1),
    secrets status exits 1 only when a required key is missing (--strict for every
    allow-listed key), and a local server that cannot start exits 2 from run and eval.
    scaffold upgrade exits 3 for a manifest without a released cli_version or an
    install-spec override without {version}, and 2 when uvx is missing or cannot fetch and
    run the prior release (all were 1); a version-locked scaffold enhance without uvx
    exits 2 (was 1).
  • Chat API changes. The SSE error event is {code, message, error_id, run_id} with
    code one of run_failed, timeout, recursion_limit, thread_busy, unavailable,
    forbidden (was the exception class name); details go to the server log (and detail only
    under APP_ENV=dev). /chat metadata outside the caps is refused with 422 (was silently
    dropped). Under langgraph-server thread ids must be UUIDs.
  • .github/agent.env is data, not shell. The workflows accept only IMAGE_REPOSITORY,
    RELEASE_NAME, CHART_PATH, RUNTIME, CD and GRAPH_AGENTS_CLI_SPEC, and refuse
    anything else.
  • A run that reaches the step limit ends with a reply, status step_limit.
    RECURSION_LIMIT defaults to 50 (was 25): two steps to answer plus two per sequential tool
    call, so 24 calls. A run that reaches it streams a final reply saying so and ends with
    message.end "status": "step_limit" (was an error event recursion_limit, now sent only
    when the reply cannot be written); its work stays in the thread. Clients that treat any
    status other than ok as a failure should accept step_limit; the eval client counts it as
    an error turn ("message.end status step_limit").
  • Run records and statuses. A run is recorded when it starts (running) and ends ok,
    step_limit, error, timeout, cancelled or interrupted (its lease was lost, or its
    process died: reconciled about a minute after the lease expires, error_type
    ProcessLost). Dashboards keyed on the old statuses need the new ones.
  • GET /threads lists the caller's own threads by default, read-across roles included
    (they got every principal's before); ?scope=all lists every thread for a role in
    AUTH_READ_ACROSS_ROLES (403 otherwise; any other scope is 422). Rows gain owner, the
    hashed principal id.
  • One message cap for every surface. MAX_MESSAGE_CHARS (default 32000): a longer
    message is 422 on /chat and JSON-RPC -32602 over A2A 1.0 and 0.3 (A2A accepted up to
    MAX_REQUEST_BYTES before). A /chat 422 no longer echoes the submitted value (input,
    url); a too-long message is value_error (was string_too_long); text with an unpaired
    surrogate is 422 (was 500).
  • Failed tool calls reach clients as an error id. Outside APP_ENV=dev a failed call's
    tool.result result and its message in GET /threads/{id}/messages read "The tool call
    did not succeed. Reference: <error_id>." with a new error_id field; the error text
    (policy rule, limit, upstream status and reason) goes to the model only. API-policy
    refusals read "... refused by the API policy: ." (no "(api-policy.yaml)").
  • Outbound calls: stricter headers and no method override. Tool-supplied Host,
    method-override (X-HTTP-Method-Override and its underscore spelling), X-Forwarded-*,
    Forwarded, X-Original-URL, X-Rewrite-URL and hop-by-hop headers are dropped with a
    warning; a _method query parameter or top-level JSON body key raises ApiPolicyError.
  • Eval gates can change result. expect.contains and not_contains ignore case (add
    expect.case_insensitive: false for exact matching; a not_contains word now also fails
    in another case). A quality metric's pass rate is passed / scored over the cases that ran
    it (was over every planned case), so a metric declared on only some cases can now miss its
    gate. eval_config.yaml judge: accepts only provider, model and
    max_tool_result_chars, and an unknown prompt_template placeholder is exit 3 at load.
  • New projects list API_KEY in secrets.keys only under shared-bearer (the one policy
    that reads it); secrets apply, deploy and login --write-env generate it only there.
    Existing manifests keep what they list.
  • Chart: bounded shutdown and a separate metrics Secret. The chart refuses to render when
    terminationGracePeriodSeconds (30) is not above shutdown.preStopSleepSeconds (5) +
    `shutdown.dra...
Read more