graph-agents-cli 0.2.0
graph-agents-cli is now a generic CLI for building, evaluating and deploying LangGraph agents
on self-hosted Kubernetes, for any project and any domain. Nothing in the CLI, the template,
the skills or a generated project is shaped around one consumer: projects choose an auth
policy and declare the outbound APIs their tools may call. This release also closes most of
the production-readiness findings of an independent assessment of 0.1.0 (runtime guardrails,
per-user authentication, deploy safety, supply chain, release engineering); the remaining
ones are listed under "Known limitations" and "Where it is behind" in the README.
Parked medium- and low-priority issues are listed in KNOWN_ISSUES.md.
Install from the release tag (the package is not on PyPI yet):
uv tool install git+https://github.com/ss7172/graph-agents-cli@v0.2.0Breaking changes and migration
--auth-policy product-sessionis nowcustom.create --auth-policy product-session
is refused with a hint. A manifest orAUTH_POLICYthat still saysproduct-sessionis
read ascustom, with a one-line deprecation warning. In the template,
app/policies/product_session.py(ProductSessionPolicy) becameapp/policies/custom.py
(CustomPolicy). SetAUTH_POLICY=customandauth_policy: custom.- The product API policy is now a multi-API outbound policy.
product-policy.yaml
becomesapi-policy.yaml,create --product-policybecomescreate --api-policy, the
manifest blockproduct_api:becomesapi_policy: {policy_file: api-policy.yaml}, the
cookiecutter variablehas_product_policybecomeshas_api_policy, and the tool
declarationPRODUCT_CALLSbecomesAPI_CALLSwith an"api"key per entry. The file
now declares any number of APIs underapis: {<name>: ...};allowed_methodsis
required;auth: forwarded-sessionis nowauth: forward.create,scaffold enhance,
scaffold upgradeandlintstop on a project that still uses the old format and print
the migration steps (exit 3);--product-policyis refused with a rename hint; a tool
module that declaresPRODUCT_CALLSis a lint error. - Outbound API calls fail closed. Without
api-policy.yaml, or for an API the file does
not declare,get_client()raisesApiPolicyErrorand nothing is sent. 0.1.0 sent every
call unrestricted (with a warning) when no policy file existed. - The default guidance file is
AGENTS.md(wasGEMINI.md). Existing projects keep the
file name their manifest records; pass--agent-guidance-filename GEMINI.mdtocreateto
keep the old default. CLI_VERSION_PINis nowGRAPH_AGENTS_CLI_SPECin.github/agent.env(the
cookiecutter variablecli_version_pinis nowcli_install_spec). The value is a full
install spec,git+https://github.com/ss7172/graph-agents-cli@v0.2.0by default, and the
workflows runuvx --from "$GRAPH_AGENTS_CLI_SPEC" graph-agents-cli .... The workflows
refuse anagent.envthat still setsCLI_VERSION_PIN, with a rename hint.- Installation moved to a pinned git reference. The package name
graph-agents-cliwas
never published on PyPI, souv tool install graph-agents-clinever worked;setup,
update, thescaffold upgradebaseline and the generated workflows install
git+https://github.com/ss7172/graph-agents-cli@v<version>instead, and the update check
reads GitHub releases.GRAPH_AGENTS_CLI_INSTALL_SPECoverrides the source (a mirror, a
wheel); write{version}where the version goes. deployandsecrets applyoutsidedevneed an explicit env file and kube context.
They read.env.<env>(or--env-file) and never fall back to.env(exit 3 without one).
The kube context must be recorded asenvironments.<env>.contextor passed with
--context; the kubeconfig's current context is used only after a confirmation prompt, or
--yeswhen there is no terminal (exit 1 otherwise). CI jobs that deploy need--yesor
--context(the generated workflows pass both).- A live
API_KEYis never replaced implicitly. Undershared-bearer,secrets apply
anddeploykeep the key in the cluster unless the env file sets a different one and
--rotate-api-keyis passed. A generated key is written to the env file (mode 0600)
instead of being printed. helm-pushreadsDEPLOY_KUBECONFIGfrom thestagingandproductionGitHub
environments (was the repository secretKUBECONFIG), so the production reviewers gate
it. Create the environment secrets and delete the repository secret.deploy --env staging|prodis refused outside CI inhelm-pushmode even with--image
(--force-directoverrides).- Chart defaults are stricter.
image.tagdefaults to""in every environment and the
chart refuses to render without a tag (neverlatestby default); the tag must be a
quoted string. The HTTPRoute and Ingress publish onlyroute.publicPaths(/chat,
/threads,/a2a/<agent>, plusroute.devPathsunderAPP_ENV=dev) instead of every
path;/health,/readyand/metricsstay inside the cluster. Outsidedevthe app
Secret is required (secretOptional: false): pods do not start without it. - Exit codes are consistent (0 ok, 1 refused or failed gate, 2 tool failure, 3
configuration error): an unexpected crash exits 2 (was 1), running outside a project
exits 3 (was 1),runexits 2 when the agent cannot be reached or goes silent (was 1),
secrets statusexits 1 only when a required key is missing (--strictfor every
allow-listed key), and a local server that cannot start exits 2 fromrunandeval.
scaffold upgradeexits 3 for a manifest without a releasedcli_versionor an
install-spec override without{version}, and 2 whenuvxis missing or cannot fetch and
run the prior release (all were 1); a version-lockedscaffold enhancewithoutuvx
exits 2 (was 1). - Chat API changes. The SSE
errorevent is{code, message, error_id, run_id}with
codeone ofrun_failed,timeout,recursion_limit,thread_busy,unavailable,
forbidden(was the exception class name); details go to the server log (anddetailonly
underAPP_ENV=dev)./chatmetadata outside the caps is refused with 422 (was silently
dropped). Underlanggraph-serverthread ids must be UUIDs. .github/agent.envis data, not shell. The workflows accept onlyIMAGE_REPOSITORY,
RELEASE_NAME,CHART_PATH,RUNTIME,CDandGRAPH_AGENTS_CLI_SPEC, and refuse
anything else.- A run that reaches the step limit ends with a reply, status
step_limit.
RECURSION_LIMITdefaults to 50 (was 25): two steps to answer plus two per sequential tool
call, so 24 calls. A run that reaches it streams a final reply saying so and ends with
message.end"status": "step_limit"(was anerroreventrecursion_limit, now sent only
when the reply cannot be written); its work stays in the thread. Clients that treat any
status other thanokas a failure should acceptstep_limit; the eval client counts it as
an error turn ("message.end status step_limit"). - Run records and statuses. A run is recorded when it starts (
running) and endsok,
step_limit,error,timeout,cancelledorinterrupted(its lease was lost, or its
process died: reconciled about a minute after the lease expires,error_type
ProcessLost). Dashboards keyed on the old statuses need the new ones. GET /threadslists the caller's own threads by default, read-across roles included
(they got every principal's before);?scope=alllists every thread for a role in
AUTH_READ_ACROSS_ROLES(403 otherwise; any other scope is 422). Rows gainowner, the
hashed principal id.- One message cap for every surface.
MAX_MESSAGE_CHARS(default 32000): a longer
message is 422 on/chatand JSON-RPC -32602 over A2A 1.0 and 0.3 (A2A accepted up to
MAX_REQUEST_BYTESbefore). A/chat422 no longer echoes the submitted value (input,
url); a too-long message isvalue_error(wasstring_too_long); text with an unpaired
surrogate is 422 (was 500). - Failed tool calls reach clients as an error id. Outside
APP_ENV=deva failed call's
tool.resultresultand its message inGET /threads/{id}/messagesread "The tool call
did not succeed. Reference: <error_id>." with a newerror_idfield; the error text
(policy rule, limit, upstream status and reason) goes to the model only. API-policy
refusals read "... refused by the API policy: ." (no "(api-policy.yaml)"). - Outbound calls: stricter headers and no method override. Tool-supplied
Host,
method-override (X-HTTP-Method-Overrideand its underscore spelling),X-Forwarded-*,
Forwarded,X-Original-URL,X-Rewrite-URLand hop-by-hop headers are dropped with a
warning; a_methodquery parameter or top-level JSON body key raisesApiPolicyError. - Eval gates can change result.
expect.containsandnot_containsignore case (add
expect.case_insensitive: falsefor exact matching; anot_containsword now also fails
in another case). A quality metric's pass rate is passed / scored over the cases that ran
it (was over every planned case), so a metric declared on only some cases can now miss its
gate.eval_config.yamljudge:accepts onlyprovider,modeland
max_tool_result_chars, and an unknownprompt_templateplaceholder is exit 3 at load. - New projects list
API_KEYinsecrets.keysonly undershared-bearer(the one policy
that reads it);secrets apply,deployandlogin --write-envgenerate it only there.
Existing manifests keep what they list. - Chart: bounded shutdown and a separate metrics Secret. The chart refuses to render when
terminationGracePeriodSeconds(30) is not aboveshutdown.preStopSleepSeconds(5) +
shutdown.drainSeconds(20); raise it with them. The ServiceMonitor's bearer token now
comes from the Secret<release>-metrics(was the app Secret). deployrefuses more outsidedev: aCHANGE-MEvalue in the chartenv(exit 3; a
warning indev), andjwtwithout a JWKS URL or public key,AUTH_JWT_ISSUERand
AUTH_JWT_AUDIENCE(exit 3).deploy --statusand--restartwait at most--timeout
and exit 1 / 2 when the pods are not ready (they returned at once before).extension add/updatewithout a terminal never prompts: untrusted code needs-y
(exit 1 otherwise;updatekeeps the installed copy).
Upgrading a running deployment
- Roll out with
Recreate, or at one replica. Runs take a Postgres lease per thread
(thread_locks,agent_thread_locksunderlanggraph-server), which older builds do not
honour: 0.1.0 has no run lock across replicas, and pre-release 0.2.0 builds used a session
advisory lock. While old and new pods run side by side, one thread can run on both. - The new tables, their token sequence and a partial index on the runs table (built
CONCURRENTLY) are created at startup; nothing to run by hand. - With
metrics.serviceMonitor.bearerToken.enabledand the default Secret, run
graph-agents-cli secrets apply --env <env>once after upgrading the chart, so
<release>-metricsexists (infra checkshows the row). - If you lowered
terminationGracePeriodSeconds, keep it above
shutdown.preStopSleepSeconds + shutdown.drainSeconds(or lower those). - Clients: accept
message.endstatusstep_limit; list other principals' threads with
GET /threads?scope=all; let the server generate thread ids (omitthread_id) or use
UUID4s: a thread id another principal used first is theirs. - Eval datasets: review
not_containschecks (now case-insensitive) and metrics declared
on only some cases (see above). - Template files you have not edited take the new versions with
graph-agents-cli scaffold upgrade(a project made by a pre-release 0.2.0 build names that build: see "Upgrading a
project made by a pre-release 0.2.0 build" below); inapp/agent.pykeepmiddleware()(SurfaceApiErrors,
AnswerInvalidToolCallsandUntrustedToolResults, in that order) if you rewrote it,
and in tools useToolRuntime[Any].scaffold upgradenever rewritesagent.py: an
edited one needsAnswerInvalidToolCalls()added by hand (fromapp_utils.content).
Upgrading a project created with 0.1.0
-
If the project uses
product-policy.yaml, migrate it first: every project command prints
the steps (rename the file toapi-policy.yaml, move the fields under
apis: {<name>: ...}withallowed_methods, changeproduct_api:in the manifest to
api_policy: {policy_file: api-policy.yaml}, renamePRODUCT_CALLStoAPI_CALLSwith an
"api"key). -
Run
graph-agents-cli scaffold upgrade(preview with--dry-run, apply with-y). Its
authentic baseline re-renders the project with graph-agents-cli 0.1.0 from thev0.1.0
tag of this repository (commitfc3f2f9), so it updates every scaffolding file you did
not edit (about 30:app/app_utils/*.py,app/fast_api_app.py, the Dockerfile, the chart
templates andvalues.yaml,pr_checks.yaml,.github/agent.env, ...), adds the new
ones (app/policies/custom.pyamong them), removesapp/app_utils/product_client.pyand
reports your own edits as conflicts. It needsuvxand access to the repository.If the tag cannot be fetched (it is not on the remote yet, an offline mirror, a fork
without tags), name the same build in any clone that holds commitfc3f2f9:git clone https://github.com/ss7172/graph-agents-cli /tmp/gac # or your mirror graph-agents-cli scaffold upgrade --baseline-ref /tmp/gac@fc3f2f9 --dry-run graph-agents-cli scaffold upgrade --baseline-ref /tmp/gac@fc3f2f9 -y(
GRAPH_AGENTS_CLI_INSTALL_SPEC='git+file:///tmp/gac@v{version}'after
git -C /tmp/gac tag v0.1.0 fc3f2f9still works, but the override also becomes the new
.github/agent.env'sGRAPH_AGENTS_CLI_SPEC, which you then set back to
git+https://github.com/ss7172/graph-agents-cli@v0.2.0.)Do not use
--baseline currentfor a 0.1.0 project. It compares against the 0.2.0
templates, so it cannot tell your edits from 0.2.0's changes: every scaffolding file 0.2.0
changed is listed under "Will preserve" and keeps its 0.1.0 content, the new dependencies
(pyjwt,prometheus-client) are not merged, and only new files are added (not
app/policies/custom.py). After step 3 the project's tests fail to import and/chat
answers 500. If you ran it, restore the project from the backup it printed
(~/.graph-agents-cli/backups/...) or from git, and upgrade with the authentic baseline. -
scaffold upgradenever rewrites agent code or config, so port these by hand (compare
with a freshgraph-agents-cli createof the same settings):app/policies/__init__.py: replace it with the 0.2.0 registry. The old file imports
PRODUCT_SESSION, which no longer exists, so the app does not start until it is
replaced. If you implementedProductSessionPolicy, move it into
app/policies/custom.pyasCustomPolicyand deletepolicies/product_session.py.app/agent.py: importApiCallError,ApiPolicyErrorfromapp_utils.api_client
(the old file imports the removedapp_utils.product_client) and bind
recursion_limit(); an unmodified file can be replaced with the new one.app/tools/: deleteproduct_lookup.py(or port it toget_client()), rename
PRODUCT_CALLStoAPI_CALLSin every module (weather.pyincluded), and take the new
tools/__init__.py.deployment/helm/<name>/values-dev.yaml: addsecretOptional: true(without it the dev
pods wait for the app Secret); setimage.tag: ""in eachvalues-<env>.yaml(or a
quoted tag) instead oflatest; take the prodresourcesandtopologySpreadif wanted..env.example: compare with a fresh render for the new variables.
-
graph-agents-cli install,uv run pytest tests/unit tests/integrationwith
MODEL_PROVIDER=fake, andgraph-agents-cli lint.
Upgrading a project made by a pre-release 0.2.0 build
Builds made before the release share its version, so such a project's manifest says
cli_version: '0.2.0' without the cli_build record this release adds, and scaffold upgrade
answers "already at version 0.2.0" (compared by version only) with the steps below. Nothing in
the manifest has to be edited.
-
Find the commit of the build that created the project. If you do not know it, the
newest commit before the project was generated is a first candidate:
git -C <checkout> log -1 --format=%H --before='<generated_at from the manifest>'. The
build may be older than that: a checkout behind its branch, or a stale build
(uv tool install --fromreused its cached wheel of an earlier commit before this
release, see Fixed). -
Preview with that build as the baseline, then apply (a local clone reaches commits that
were never pushed). With the right build, only files you edited are listed under "Will
preserve" or as conflicts; many scaffolding files you never touched there mean the wrong
build, so try an earlier commit:graph-agents-cli scaffold upgrade --baseline-ref <checkout>@<commit> --dry-run graph-agents-cli scaffold upgrade --baseline-ref <checkout>@<commit> -y
The baseline must be a 0.2.0 build (exit 3 otherwise). Files you did not edit take the
0.2.0 versions, new ones are added, your edits are kept or reported as conflicts, and the
manifest then records this build incli_build, so later upgrades need no flag. -
Port what
scaffold upgradenever rewrites (app/agent.py,app/tools/**, the
values-<env>.yamlfiles): see "Upgrading a running deployment" step 7.
Added
- The manifest records the build that rendered the project, as
cli_build: its id (what
graph-agents-cli --versionprints:0.2.0for the release,0.2.0+g<commit>between
releases), its full commit andtemplate_digest, a digest of what that build renders for
the recorded settings (null whencreateseeded a policy or used a local or remote
template).createwrites it (keeping the manifest's comments),scaffold upgradeand a
settings change withscaffold enhancerewrite it, andinfoshows it
(Scaffolded with: 0.2.0 (build ...)).scaffold upgradeuses it to pick the old
snapshot's build: a build between releases is rebuilt from its commit; at the running
version a project is up to date only when the build, or what it renders, is the same. A
recorded build with uncommitted changes, or a build between releases while
GRAPH_AGENTS_CLI_INSTALL_SPECis set (its{version}names releases only), stops the
upgrade with the ways out (exit 3). scaffold upgrade --baseline-ref REFnames the build that created a project when the
manifest cannot: a commit or tag of the repository,<clone>@<commit>for a local clone
(looked up there first), a path to a checkout or wheel (rebuilt withuvx --refresh-package), or a full install spec. The baseline must render the manifest's
cli_version(exit 3 otherwise); a different recorded commit is a warning, and so is a
baseline under which most template files would keep their current content (the sign of a
later build than the one that created the project) or that renders the same files as the
running build.jwtauth policy: per-user principals from a verified OIDC/JWT bearer token. JWKS URL
(AUTH_JWT_JWKS_URL, cached forAUTH_JWT_JWKS_CACHE_S, one rate-limited refetch on an
unknown key id, stale-while-revalidate, a bounded grace when the issuer is down) or one PEM
key (AUTH_JWT_PUBLIC_KEY); issuer and audience (required outsideAPP_ENV=dev),
exp/nbf/iatwithAUTH_JWT_LEEWAY_S; an algorithm allow-list (AUTH_JWT_ALGORITHMS,
defaultRS256,ES256; nevernone; HS256/384/512 only withAUTH_JWT_ALLOW_HS=trueand a
32-byteAUTH_JWT_SECRET);AUTH_JWT_PRINCIPAL_CLAIMandAUTH_JWT_ROLES_CLAIM(dotted
paths); https JWKS outside dev unlessAUTH_JWT_JWKS_ALLOW_HTTP=true. RFC 6750 challenges;
nothing from the token is logged.customauth policy as a documented, fail-closed interface (authenticate,
authorize, optionalstartup_problems()), andPrincipal.public_attributes(): secrets
live only underattributes["credentials"]and are never persisted, logged or traced.- Startup fails closed: an unknown
AUTH_POLICYnever starts; a misconfigured policy stops
the process outsideAPP_ENV=dev. AUTH_ADMIN_ROLES: underlanggraph-server, only these roles may create, update or delete
assistants and crons or write the store (default: nobody); reads are open to authenticated
principals; every other native-API action is denied by default.api-policy.yaml(see Breaking changes): strict schema shared byte-for-byte by
create,lintand the runtime client (app_utils/api_client.py),auth: none | bearer | forward,allowed_operations/denied_operations(denials win and fail closed),
OpenAPI validation inlint, timeouts, an enforced pagination cap, path-prefix-safe URL
joins, no redirects. There is no default access level: every API lists its methods
explicitly.create --api-policyvalidates the file first, copies the OpenAPI specs it
references and renders an example tool making the first operation the policy's first API
allows, whatever its method (abodyargument for POST, PUT and PATCH); bearer tokens join
secrets.keys;auth: forwardis refused underlanggraph-server. The runtime client
sends every allowed method (request(),get,head,post,put,patch,delete,
options) with JSON bodies, query parameters and headers.graph-agents-cli api: the policy belongs to the project and evolves with the agent.
api add NAME --base-url-env ENV --auth none|bearer|forward --access read-only|read-write|custom(--accessis required: read-only = GET, HEAD; read-write =
GET, HEAD, POST, PUT, PATCH, DELETE; custom =--methods),api access,api allow/
api deny(by operationId, filled in from the API's OpenAPI spec when it has one; by
--method/--path; or by both, pinning all three),api revoke,api limits,api remove,api show [--json]andapi check(same aslint --policy-only). Every change
validates the result, keeps comments and key order, prints a unified diff of each file it
touches (the policy, the manifest'sapi_policyandsecrets.keys,.env.example, the
chart'svalues.yaml), writes atomically, says whether it widens or narrows access and how
the tools' declared calls are affected;--dry-runprints the diff only; exit 3 on an
invalid result, outside a project, or when the edit would also change another API that
repeats the edited one through a YAML alias.lintprints everygraph-agents-cli api
command a refused call needs (the method, a denial, an allow-list entry pinning the call's
method and path), andlintandapi checkexit 3 on an invalidapi-policy.yaml(a
configuration error, not a refused call).- Outbound call limits: optional per-API
limits: {max_calls_per_run, rate_per_minute}.
max_calls_per_runcounts the calls to that API within one agent run (the LangGraph run
id, else the request's);rate_per_minuteis a token bucket per process (per replica).
A call over a limit is refused before it is sent, with a reason the model can read;
a run's counters are dropped when a/chator A2A run ends, and otherwise (LangGraph
Server runs included) after an hour without a call. - Human approval of calls: an API's optional
approvalblock (required_for: {methods, operations},approvers: [requester | "role:<name>"],timeout_s30-86400, default 900)
makes those calls wait for a person. Approval never widens access: a gated call must still be
allowed, denials still win, and anapprovalkey on an operation entry is refused with a
pointer toapproval.required_for.operations(whose entries hold like denials, whatever
label a call gives). The run pauses before sending a gated call:/chatends with
message.endstatusawaiting_approvaland the call (API, method, full path, query, body,
operation id, reason, approvers,expires_at);GET /threads/{thread_id}/approvalslists a
thread's approvals andGET /approvalsthose the caller may see across threads (its own, the
ones naming one of its roles);POST /threads/{thread_id}/approvals/{approval_id}with{"decision": "approve"|"reject", "comment"}decides one (403 for a non-approver, 404, 409 once decided,
410 once expired) and streams the resumed run; a new/chatmessage on a paused thread gets
409approval_pending; an A2A task goesinput-requiredand resumes with a data part
carrying the decision.requesteris the principal who started the run,role:<name>any
other principal holding the role. An approved call is sent exactly as shown, once, and only
while the policy still allows it and gates it with the same approvers; a rejected or expired
one never, whatever the policy says about gating it by then (a decision is bound to its
call), and a call still waiting when another decision resumes the run waits on for its own.
Approvals are kept in anapprovalstable beside the checkpoints (agent_approvalsunder
langgraph-server; under the locallanggraph dev, in.langgraph_api/agent_approvals.json
beside its threads, so both survive a restart or a hot reload); deleting a thread deletes
them. - Approval rules: other approvers for other calls of one API.
approvalmay also be a
non-empty list of rules of the same shape (each with its ownrequired_for,approvers,
timeout_s), so one API can have the requester confirm updates and cancellations while a
role:adminapproves new orders, without declaring the API twice. A call is gated by the
first rule in file order whoserequired_forcovers it, with that rule's approvers and
expiry; later rules that also cover it do not apply to it. The approvers recorded with a
pending approval, who decide it, are those of the rule that gated the call when it paused,
and an approved call is sent only while the rule that gates it then names the same
approvers. A single mapping keeps its meaning; every rule of a list is validated in full
(errors nameapproval[N]), fail-closed matching and "approval never widens access" hold
per rule, and an operation-levelapprovalkey stays invalid. A call that the first rule
covers only because it names no operation id (an entry byoperationIdalone), and that a
later rule with other approvers also covers, is refused: it could be either rule's call, so
neither rule's approvers decide it (lintandapi checkreport such declared calls, and
api approvalnotes the rule). The CLI and the runtime share the rules byte for byte. graph-agents-cli api approval NAME[--methods M,...|none] [--operations OP,...|none] [--approvers requester,role:NAME] [--timeout-s N] [--add-rule | --rule N] [--remove] [--dry-run], with the otherapicommands' validate, diff and atomic-write rules: each
option replaces that part of the rule, operations are pinned to their method and path from
the API's OpenAPI spec, and the command says when a change loosens the gate (a reviewed
change), which declared calls become gated and which get other approvers.--add-rule
appends a rule (turning a single block into a list, comments kept),--rule Nchanges or
removes rule N, and a command on a list of several rules without either is refused.
lint,api checkandapi show(--json:approvalandapproval_rulesper API;
approvalper call with itsrule,rule_indexandalso_covered_by; agatedcount)
list which declared calls wait for whose approval and by which rule, count calls several
rules cover, and note a rule that never applies.graph-agents-cli approvals list|approve|rejectfor the project's local server or a
deployed agent (--url), with the credentialsrunsends (GRAPH_AGENTS_CLI_API_KEY,
--header,--cookie):listshows a thread's approvals, or every one the caller may see
(so arole:approver needs no thread id);approve/rejectshow the call, send only the decision and the comment, and
stream the resumed run.runprints a paused call in full and, on a terminal when the
requester is an approver, asksApprove? [y/N]and continues; otherwise it prints the
decision commands and exits 0 with an "Awaiting approval" line, keeping a one-off local
server with the in-memory checkpointer running so the paused run survives. With no local
server running,approvalsstarts a temporary one where the paused run and its approval
outlive their server (fastapiwith the postgres checkpointer, andlanggraph-server).- Eval approvals: a dataset case declares how a human would decide each gated call it
reaches ("approvals": [{"decision": "approve"|"reject", "match": {"operation_id": ...} | {"method": ..., "path": ...}}], optionally withapi);eval generatedecides each gate
per the first matching instruction and continues the run (a gate no instruction matches is a
case error, and is rejected, as is one whose decision the server refused; one the eval may
not reject goes with the case's thread, which the eval identity deletes, so no approval is
left pending:approvals[].cleanupin the trace), traces record every gate, andexpect.approvals(gated,approved,
rejected) andexpect.no_approvalscheck them. A gate that listsrequesteris decided
as the eval identity, any other asGRAPH_AGENTS_CLI_APPROVER_API_KEYwhen set. infra checkreports where pending approvals are kept (theapprovalstable of the app
database) and warns whenCHECKPOINTER=memorywould lose paused runs on a restart.- Endpoints:
GET /ready(readiness: the database is set up and answers within 2 s),
GET /metrics(Prometheus: request count and latency, runs by status, active runs, run
duration, tokens, database up; optionalMETRICS_TOKEN),GET /threads(the caller's
threads;?scope=allfor read-across roles) andDELETE /threads/{thread_id}. - Runtime guardrails: one run per thread (409
{"code": "thread_busy"}, with a Postgres
lease across replicas),RUN_TIMEOUT_S,MODEL_TIMEOUT_S,MODEL_MAX_RETRIES,
RECURSION_LIMIT(50),MAX_REQUEST_BYTES(413),MAX_MESSAGE_CHARS,MAX_METADATA_KEYSand
MAX_METADATA_VALUE_CHARS(422),SSE_HEARTBEAT_S, a client disconnect cancels the run,
and a stopped run answers its open tool calls so the thread stays usable. RETENTION_DAYS: an hourly best-effort purge of threads (checkpoints and run records) idle
longer than N days.- Structured JSON logging (
LOG_FORMAT,LOG_LEVEL) with request id (X-Request-IDon
every response), run id, thread id and a hashed principal; optionalPRINCIPAL_HASH_SALT
keys the hash (HMAC-SHA256). Client-facing errors carry anerror_id; details stay in the
log. CORS_ALLOW_ORIGINS(empty: no CORS),DB_POOL_MIN_SIZE/DB_POOL_MAX_SIZE,
AUTH_FORWARD_HEADERS,A2A_TASK_TTL_S,API_POLICY_PATH.deploy:--context,--yes,--timeout(default 5m),--atomic/--no-atomic(default
atomic),--rotate-api-key; the resolved kube context and API server are printed before
anything happens; a pre-deploy check that the Secret holds every required key; pod
diagnostics (states, warning events, logs) when a rollout fails, then a rollback of this
run's revision only; a refusal while another helm operation holds the release; workstation
images from a dirty tree are tagged<sha>-dirty-<timestamp>.secrets apply:--context,--yes,--rotate-api-key; creates the namespace when it is
missing.secrets status:--context,--strict, exit codes usable as a gate.infra check: rows for unreplacedCHANGE-MEplaceholders (registry, chart image, chart
env, CODEOWNERS, Argo CDrepoURL), missing required Secret keys, and, forhelm-push,
whetherDEPLOY_KUBECONFIGexists as an environment secret and no repository-level
kubeconfig secret exists.- Chart:
metrics.serviceMonitor.bearerTokenmakes the ServiceMonitor sendMETRICS_TOKEN
(from the Secret<release>-metricsby default, whichsecrets applyanddeploywrite
with that key alone), so a token-protected/metricscan be scraped. run --portandGRAPH_AGENTS_CLI_RUN_PORT; a port preflight forrunandplayground
(exit 3 when the port is taken);GRAPH_AGENTS_CLI_DEBUG=1shows the traceback behind a
one-line error.scaffold enhance --runtime/--model-providerreconciles everything the change affects
(model default,secrets.keys,.env.example, chart values merged key by key around your
edits,.github/agent.env) and ends with a "Left for you" list; steps marked(required)
make it exit 1, and an edited Dockerfile gets the new version beside it as
Dockerfile.new.- Chart: readiness on
/ready, liveness and startup on/health; default requests and
limits (100m / 256Mi, 1Gi limit; 250m / 512Mi requests in prod); read-only root filesystem with a/tmp
emptyDir, uid/gid 1000, seccompRuntimeDefault, all capabilities dropped, no service
account token; optional NetworkPolicy, ServiceMonitor and scrape annotations (off by
default); soft topology spread in prod; a chart-managed dev Postgres password that
survives upgrades; subcharts pinned to exact versions and their images by digest. - Generated workflows:
pr_checksruns the tests on the fake model and the eval gate on the
real provider when its key secret exists (with a warning when the gate runs on the fake
model); the helm-push jobs resolve one kube context, pass--contextand--yes, and
verify the rollout (/health,/ready);stagingcan be dispatched by hand frommain
and does not re-trigger itself; argocd staging PRs supersede older ones;
GRAPH_AGENTS_CLI_DISABLE_OVERRIDES=1in every CI and CD job; one.github/CODEOWNERSfor
every project coveringdeployment/,.github/,api-policy.yaml,tests/eval/,
extensions and the manifest. - Release engineering for the CLI itself: this changelog,
.github/workflows/ci.yml(ruff
and the fast suite on every pull request and push tomain; the end-to-end suite nightly
and on demand) and.github/workflows/release.yml(a tagvX.Y.Zbuilds the sdist and
wheel and creates a GitHub Release; PyPI trusted publishing is opt-in). graph-agents-cli auth dev-token --sub USER [--roles R,...] [--ttl 12h]: a JWT for
local runs of ajwtproject (APP_ENV=devonly), printed alone on stdout for
export GRAPH_AGENTS_CLI_API_KEY="$(graph-agents-cli auth dev-token --sub alice)". The
first call creates a dev RSA key in.graph-agents-cli/dev-jwt/(0600, git ignored; safe
when several runs start at once) and fills blankAUTH_JWT_PUBLIC_KEY,AUTH_JWT_ISSUER
andAUTH_JWT_AUDIENCEin.env. Refused (exit 3) for other policies, outside dev, or
with a JWKS URL or another key configured.logingains the checksjwt_key,
jwt_token,jwt_claimsandenv_file.- History repair: a thread left mid tool call (a timeout, a disconnect, a crash, an OOM
kill, a database outage) is repaired at the start of the next run on both runtimes: each
open call gets an error result right after it, and a result written after a later message
is moved back. Calls and results are paired turn by turn, so tool-call ids that repeat
across turns (call_0in every message) never make a healthy thread look damaged, and a
healthy thread is never rewritten. - Run leases: one run per thread across replicas is a Postgres lease with a 30 s expiry,
renewed every 5 s, with a fencing token checked before every checkpoint write. A frozen or
partitioned replica frees its threads after 30 s (a session advisory lock held them for
about 2 hours); a database restart or a killed session no longer lets a second replica run
the thread. A run that cannot renew its lease stops (interrupted) before it writes. - Database outages: connections default to
connect_timeout=5and TCP keepalives (the
DSN wins); requests answer 503 "Database unavailable. Reference: " within 5 s (2 s once
the app knows) with one WARNING line and no traceback, on every route; the app starts and
stays alive while Postgres is unreachable (/health200,/ready503) and becomes ready
once it answers. New metricagent_database_up. The pool checkout timeout is 5 s (was 10). - A startup WARNING when an API's
limits.max_calls_per_runcannot be reached within
RECURSION_LIMIT(it names the value needed). - Untrusted tool output:
app_utils.content.UntrustedToolResults, wired into the
generated agent next toSurfaceApiErrors(agent.middleware()), fences every tool
result the model reads in<tool_output ... trust="untrusted">tags on both runtimes,
whatever the result's text looks like (tags inside it are renamed, so it cannot close the
fence or forge one); the defaultSYSTEM_PROMPTsays tool output is data, never
instructions. Helpers for tools inapp_utils.api_client:require_user_mentioned,
require_owner,current_caller,latest_user_message. Content blocks are fenced as one
text (a tag split across two blocks is renamed too), other non-media blocks are read as
JSON text, and look-alikes of the tag (full-width brackets, zero-width characters, HTML
entities) are renamed as well. - Tool calls with arguments that are not valid JSON:
AnswerInvalidToolCalls(in
app_utils.content), wired intoagent.middleware(), answers each one with an error
result saying so and asks the model again in the same step (at most twice; it adds no
graph step, soRECURSION_LIMITcounts the same)./chatstreams such a call as
tool.call(withargs: {}) and an errortool.result, and the thread history lists it. - A2A:
A2A_DESCRIPTIONsets the card's description (and its chat skill's) and is in the
chart values;AGENT_VERSIONsets the card version.SendMessagereturns the reply as one
text part; streamed replies mark the last chunklastChunk, and the stored task holds one
part. A message with no text, an empty text part, a non-user role or over
MAX_MESSAGE_CHARSis -32602, and an A2A 0.3 request that fails the SDK's validation
(a JSON-escaped method name, an unpaired surrogate, a missingmessageId) is answered
-32602 or -32600 naming the fields, never the values, with no traceback. A 0.3 request
that names an unknown or deleted task (tasks/get,tasks/cancel,tasks/resubscribe)
or push notifications gets the code A2A 1.0 answers with (-32001, -32003, ...), logged at
INFO, instead of -32603 with a traceback. /chatwithout athread_idstarts a thread with a server-generated UUID4;DELETE /threads/{id}and the retention purge (both runtimes) drop the thread's A2A tasks.- eval: judges of a multi-turn case see every earlier turn (user message, each tool call
with its result, the agent's reply) and a new{transcript}placeholder; a tool result is
cut for the judge atjudge.max_tool_result_chars(default 50000,nullnever; was a
silent 2000) with a[TRUNCATED ...]marker, a warning andjudge_notes;
expect.case_insensitive;expect.scope: final_turn | all_turns; results add
quality.<m>.scored,passedandstatus,fake_modelandwarnings;eval grade
warns when the agent or the judge ran on the fake model ("gate met ... (fake model:
plumbing check only, not a quality signal)"), when a--urlrun's project settings name
the fake model, and whenall_turnschecks had to read the final turn only;eval generate --url/eval run --urlwarn before the first case that tools run for real
there, naming the write methodsapi-policy.yamlallows; trace files recordtargetand
model_provider, andmodelonly for the local server (nullfor--url: the agent
there does not report its model, and the project's settings need not be what runs
there);eval metric listshows the check modifiers. - deploy: a failed rollout that is rolled back (or a first install that is uninstalled)
puts the app Secret and<release>-metricsback to their values from before the run, keys
the run removed included, or deletes a Secret the run created; a Secret changed by someone
else meanwhile is left alone (and the metrics Secret with it, so both keep the same
METRICS_TOKEN), and the error says what happened.deploy --status(default--timeout
60s) prints replicas, image, helm revision and each pod's state; failed-deploy diagnostics
show only this release's warning events since the run started; a redeploy of the running
image says what will happen and advises--restartwhen only the Secret changed;
--dry-runreads the live Secret and refuses a missing required key like the real run.
An external DSN withoutsslmode=require|verify-ca|verify-fullis a warning outside dev. - Chart: a
preStoppause (shutdown.preStopSleepSeconds, 5) and a bounded drain
(shutdown.drainSeconds, 20, asUVICORN_TIMEOUT_GRACEFUL_SHUTDOWNor
BG_JOB_SHUTDOWN_GRACE_PERIOD_SECS), so in-flight requests end before the kill (fastapi:
runs endcancelledand release their leases); the dev Postgres stops in fast mode; a
NetworkPolicy example (examples/networkpolicy.yaml, not packaged) that values-staging and
values-prod point to. infra checkrows:auth: jwt settings,metrics token secret <name>-metrics,database tls (<KEY>); a disabled gateway is one skip row.run: after anerrorevent the footer prints the run, the thread and the resume command;
a dropped stream says the run started and was interrupted (and whether the thread
survived: a stopped local server with an in-memory checkpointer loses it);run -v
prints one compact line per event.- README: an "External database" section (a least-privileged role,
sslmode=verify-full,
the CA mount).
Changed
-
runexits 0 when the run is left awaiting an approval (the "Awaiting approval" line says
so) and 1 when the server refuses a decision; streamed agent text and tool output are
printed with terminal control characters escaped (line breaks and tabs kept). -
setupinstalls the skills from this repository at the tag of the running release
(https://github.com/ss7172/graph-agents-cli#v<version>; the default branch only for a
development build), so they cannot drift from the CLI when the default branch moves on;
updatemoves them to the tag of the release it installs. -
scaffold enhanceno longer lists--api-policyin its help (it refuses the flag) and
points tograph-agents-cli apifor changing the policy. -
The printed "Get Started" after
createincludescp .env.example .envand
graph-agents-cli login --write-env, so a firsteval rundoes not fail with 503. -
With no
--registryand no gitoriginremote,createstill records the placeholder
ghcr.io/CHANGE-ME; its hints, the generated README,deploy,buildandinfra check
now namegraph-agents-cli scaffold enhance --registry <host>/<org>, which sets the
registry everywhere it is read (the manifest, the chart values,.github/agent.env). -
A
TRACE_CAPTUREother thanmetadataorfullstops the app at startup, like the other
settings that do not parse (0.1.0 read it asmetadata); so does anA2A_TASK_TTL_Sthat
is not a whole number >= 0. -
scaffold upgrade --baseline currentlabels its "Will preserve" list as files that differ
from the current template, and says that files you did not edit keep their old content and
that dependency changes are not merged; when the authentic baseline fails, the hint that
suggests--baseline currentwarns about this too. -
login --write-envfills a blankKEY=line in place (no duplicate lines), keeps.env
at mode 0600 and writes it atomically. -
extension addaccepts a local path without thelocal@prefix;extension update
reports "Already up to date". -
lintchecks every*.pyunder<agent>/tools/, subpackages included, and reports any
API_CALLSit cannot read as one literal (+=,.append(), a conditional assignment). -
Backups under
~/.graph-agents-cli/backupsare private (0700,.env*files 0600) and only
the newest 5 per project are kept. -
Images: the fastapi image is multi-stage on
python:3.12.14-slim-bookwormwith a pinned
uv and no uv in the final image; the server image islangchain/langgraph-api:0.14.4-py3.12
(checked againstuv.lockat build time) with its unauthenticated meta routes disabled;
both run as uid/gid 1000 and work with a read-only root filesystem. Both make
api-policy.yamlreadable by that user whatever its mode in the working tree. -
Client
/chatmetadata is kept in the run record only: never written into checkpoints,
and exported to traces only underTRACE_CAPTURE=full. -
Log warnings print as
Warning: ...instead ofWARNING:root:.... -
Logs (both runtimes): access lines keep the path and drop the query string (under
langgraph-serverthe server's access lines lose theirquery_stringfield);httpx,
httpcore,httpx2andhttpcore2log at WARNING only (their INFO lines carry full
outbound URLs), andapi_clientlogs each outbound call asapi call done: <api> <METHOD> <operation|template> -> <status> (<ms> ms); Python warnings are records (JSON under
fastapi) with pydantic'sinput_valueredacted; a failed tool call logs one WARNING with
its error id, tool name and error type. -
.envis applied when the app module is imported (below the process environment), so the
A2A card's auth scheme,A2A_NAME, the dev-only/docsandCORS_ALLOW_ORIGINSfollow it
underuvicorntoo;PYTHON_DOTENV_DISABLEDswitches it off. -
The fake model calls whichever bound tool the request mentions (not only
get_weather)
and echoes a tool's own text (without the untrusted-data fence). Generated projects' tests
depend on none of the project's tools, its.envor the developer's shell: a
tests/conftest.pystrips app settings, server tests serve a graph without the project's
tools (@pytest.mark.project_graphopts out) and use test-only tools through
use_test_tools. -
runandevalhelp lead withGRAPH_AGENTS_CLI_API_KEYfor bearer credentials (argv is
visible to other local users); 401 hints depend on the project's policy;eval generate
prints the same hint. -
api show,api checkandlinttables grow to 250 columns in piped output instead of
cutting cells;api add --openapiprints the policy diff first and summarises the copied
spec;api allowof an entry that allows nothing yet says so. -
Deploy mode labels say "local cluster" (was "dev cluster"); generated projects git-ignore
deployment/helm/*/Chart.lock. -
login --write-envandauth dev-tokennever assign a key.envalready sets, and
concurrent writers take turns (a lock in.graph-agents-cli/);login --write-envleaves
.envat 0600.
Deprecated
- The hidden
run/evaloption--session-token: pass--header 'X-Session-Token: ...'.
It prints a warning and will be removed.
Removed
app/app_utils/product_client.pyandtools/product_lookup.pyfrom the template (replaced
byapi_client.pyandtools/example_api.py).
Fixed
scaffold upgradesaid "already at version 0.2.0" for a project made by an earlier build of
the same version, because it compared version strings only. The manifest now records the
build (cli_build, see Added) andupgradecompares it; a manifest without it is compared
by version, with a message that says how to name the build (--baseline-ref). The fallback
for a missing release tag no longer needs a tag in a clone or a relabelled manifest:
--baseline-ref <clone>@<commit>names any build.uv tool install --from <checkout> graph-agents-clicould silently reinstall uv's cached
wheel of an earlier commit: uv keyed its build cache onpyproject.tomlalone, whose
version does not change between releases.pyproject.tomlnow sets[tool.uv] cache-keys
on the commit, the tags and every file undersrc/, so a moved or edited checkout is
rebuilt (a checkout at a commit older than this fix still needs--reinstall).
--versionnames the build:0.2.0for the release,0.2.0+g<commit>for any other
commit and0.2.0+g<commit>.dirtywith uncommitted changes;infoshows the full commit
(info --json:cli_build). Every wheel and sdist built from a git checkout records the
commit ingraph_agents_cli/_build_info.json(hatch_build.py). CONTRIBUTING.md
describes installing a build from a checkout.- The
pr_checksworkflow wrote comment lines ofagent.envintoGITHUB_ENVand failed. deploycreated no namespace before applying the Secret on a first deploy.- Server runtime: "thread not found" is 404, not 503; run records are durable in Postgres.
- The schema setup races between replicas (now under an advisory lock) and the pool did not
recover after a Postgres restart (connections are health-checked). - Thread ownership is claimed atomically; thread ids are validated.
- The argocd staging workflow could re-trigger itself on its own values commit.
- Local-load (kind, k3d, minikube, k3s) was chosen from the kube context's name alone; it is
now decided from the cluster's nodes, confirmed with the kind, k3d or minikube listing. run,eval generateandplaygroundleft their local server running after SIGTERM or
SIGHUP.scaffold enhance --runtime/--model-providerleft chart values,.env.exampleand
secrets.keyson the old settings.- A thread whose tool call was cut short failed every later turn with a provider 400; the
langgraph-server repair now sends a remove-all update the real server accepts. - A tool call whose arguments were not valid JSON (
{'query': 'SF'}, a trailing comma,
query=SF, which OpenAI and OpenAI-compatible models can return) ended the run with an
empty reply and failed every later turn of the thread with a provider 400 on both
runtimes: LangChain sends such a call back as a call, and nothing answered it. The agent
now answers it in the run, and the history repair treats it as a call (a thread an older
version left broken is repaired by its next turn). - The history repair dropped a valid tool result when one assistant message repeated a
tool-call id (or gave its parallel calls empty ids), so the next turn failed with a
provider 400; results are now matched per call, not per id. - A2A: a 0.3 request naming an unknown or deleted task was answered -32603 with an ERROR
traceback, and a cancel or subscription naming one (either protocol version) left two
event-queue tasks of the SDK running, logged later as ERROR "Task was destroyed but it is
pending!". Any authenticated caller could write those records. - The untrusted-output fence could be closed from a tool result made of content blocks (a
tag split across two text blocks, which the provider joins) or with a look-alike tag. - A database outage made requests hang for about a minute with tracebacks, and a replica
started during an outage exited (crash loop);/readytook over a minute to recover. - Under
langgraph-servera run that reached the step limit ended with an error instead of
its reply (the server marks the run done after sending the error; the reply now waits
for it briefly). - A2A replies were split into one part per streamed token.
run's drop and timeout messages offered to resume threads a stopped in-memory server had
lost, and "the answer above is incomplete" when nothing had been shown.sslmode = verify-full(spaces around=, which libpq accepts) was reported as no TLS.- Under
langgraph dev(local runs of alanggraph-serverproject) every tool call through
api_clientfailed withBlockingError(the policy cache asked for the working directory
inside the event loop).
Security
- Human approval of writes moved into this release (it was planned for the next one): an
instruction planted in upstream data (a customer's order note) made a staff user's agent
cancel another customer's order and copy their data, a confused deputy the API policy alone
cannot stop because it decides which endpoints a tool may call, not on whose behalf. The
approvalblock (see Added) makes a person approve each gated call before it is sent,
bound to exactly that request and single-use. It is a per-API choice, never on by default.
runandapprovalsprint the call in full with every control, format and separator
character escaped, so nothing the model produced can hide, reorder or fake the call being
approved; streamed agent text and tool output can no longer send terminal control sequences
(colour, conceal, cursor moves); printed decision commands shell-quote the ids the server
sent (an id starting with-goes after--), and ids reach the approval routes as single
URL path segments (.and..included). A decision over A2A needs the auth policy's
approval.decideaction, as the HTTP route does; a native LangGraph Server run on a thread
whose approval is pending is refused (409), as/chatis; approvals a run left behind when
it failed or was cancelled are expired rather than blocking the thread. The API client
refuses a;in a request path (also percent-encoded): servers that strip path parameters
would route/orders/7/cancel;xto/orders/7/cancelpast a gate or a denial. It also
refuses a segment with a control character, or with whitespace at either end or next to a
dot, also percent-encoded (cancel%20,cancel%20.json,cancel%20%2e,7%00), which
servers that trim segments, trim the name before a format suffix, strip trailing dots and
spaces, or end a path at a NUL route to the gated or denied endpoint;lintrefuses the
same in declared paths, and also an encoded slash, backslash,;or dot segment
(cancel%2F,cancel%5C,cancel%3B,%2e%2e), which the client never sends.
Denials and gates also cover a literal segment's dot-suffixed spellings
(/orders/7/cancel.json,cancel.,cancel%2e), which servers that route format suffixes
or drop a trailing dot send to the gated or denied endpoint;lintand the runtime share
the rule. A decision is bound to the call it was taken for: a policy that changes while a
call waits (a new image, or a typo that un-gates it) can no longer turn a rejected or
expired approval into a send, nor let an approval outrun a later denial, narrower
allowed_methodsorallowed_operations, or a removed or changed gate. The approvals table
binds it to the tool call as well (a newmessage_idcolumn besidetool_call_id): a tool
call that runs again without a decision (a LangGraph Server run continued without input or
replayed from a checkpoint through the native API, or a copied thread) no longer sends a
rejected or expired call once its gate is removed, nor an approved call a second time; the
server's auth handler refuses such runs (403) on a thread that has approvals or waits on a
gated call, and refuses to copy a thread that has approvals. A waiting call that a later denial or narrower policy refuses when another
decision resumes the run has its approval expired, so the thread takes new messages again.
Underlanggraph dev(the local server ofrun,playgroundandevalfor
langgraph-server) the approvals were kept in memory while the server keeps its threads
across a restart or a hot reload (a code change): afterwards a run continued without input
or replayed from a checkpoint found no record and, once the gate was removed, sent a
pending, rejected or already sent call. They are now kept beside the threads, in
.langgraph_api/agent_approvals.json(mode 0600, written before each change takes effect;
a file that cannot be read stops the startup, and while one cannot be written nothing is
sent), and.langgraph_api/is in the scaffold's.gitignoreand.dockerignore.
eval generaterejects the gates it does not decide, and deletes the case's thread (with
its approvals) when it may not reject one (arole:gate without an approver credential,
or with one that may not decide it), so an unattended run leaves no approval behind for
someone to approve; only a gate whose thread cannot be deleted either stays pending (until
it expires, unless an approver approves or rejects it), named in the case error. run --mode a2ano longer follows an agent card to another origin: A2A clients dial the URL
the card advertises, and a staleAPP_URLorPORT(the template's.envsetsPORT=8000,
whichlanggraph devloads over the port it was given) sent the message and its bearer
credential to whatever listened there. A card naming another scheme, host or port is refused
with nothing sent, and the local server advertises the address it listens on (APP_URL,
unless.envsets one).setup,update, thescaffold upgradebaseline (throughuvx) and the documented
install no longer use the unpublished PyPI namegraph-agents-cli: whoever registered it
would have had their code run on users' machines. Everything installs from a pinned git tag
of this repository.- The outbound API policy cannot be widened by accident: without a policy every call is
refused; unknown and repeated YAML keys are errors; a denial pinning a path refuses every
call to it whateveroperation_idthe call gives (a relabelled or misspelt call no longer
reaches a denied endpoint), and a call that leaves out what a denial knows the operation by
is refused by it; with an OpenAPI spec,lintrefuses a declaredoperation_idthe spec
does not give that method and path;/admin/1/,/ADMIN/1and/%61dmin/1no longer get
past a denial of/admin/{x}; the page-size cap holds for every spelling of the
parameter;lintreads tool subpackages and reportsAPI_CALLSit cannot read instead of
trusting it. - A2A tasks are private to the principal that created them (another principal's task reads
as not found); the in-memory task store evicts tasks afterA2A_TASK_TTL_S. - LangGraph Server native API: assistants, crons and store writes need
AUTH_ADMIN_ROLES;
a default-deny handler covers every other resource; thread owners cannot transfer a thread
or set its tenant; read-across roles cannot copy another principal's thread; a 401 carries
the policy's challenge and a policy 503 stays a 503. - LangGraph Server: the run context carries only
public_attributes()(credentials were
persisted before), and/chatrun metadata (so traces and checkpoint metadata) carries the
hashed principal id instead of the raw one. - The server image disables LangGraph Server's unauthenticated
/docs,/openapi.json,
/infoand/metrics;/metricscan requireMETRICS_TOKEN. - Secrets are applied with server-side apply (no
last-applied-configurationannotation
holding the values; an old one is removed), and a generatedAPI_KEYis never printed. .github/agent.envcan no longer inject environment variables (BASH_ENV,PS4,
SHELLOPTS, ...) into workflow steps;GRAPH_AGENTS_CLI_INSTALL_SPECwith control
characters or stray whitespace is refused.helm-pushkubeconfig moved to environment secrets behind the production gate and is
removed from the runner after the job.- Workstation
helm-pushdeploys to staging and prod cannot bypass CI with--image. - A delete racing a chat turn can no longer leave ownerless checkpoints another principal
could claim; the retention purge re-checks idleness under the thread lock. NaN/Infinityin a request body gives 422 instead of a 500 with a traceback.- The chart runs pods as non-root with a read-only root filesystem, dropped capabilities
and no service-account token, and keeps probes and metrics off the public route. APP_ENVcounts as dev (dev-only pages, the errordetail, an optional jwt issuer and
audience, the chart's dev paths) only when it is exactlydev; 0.1.0 also accepted any
case and surrounding whitespace.- Prompt injection through tool results is reduced: results are fenced as untrusted data
(see Added) and write tools haverequire_user_mentioned/require_owner. It is not
prevented; see "Known limitations" and the langgraph-code skill (section 2a). - Query strings (a token a client put in the URL) and outbound URLs with their values stay
out of the logs on both runtimes;repr()ofPrincipaland of the run context no longer
shows forwarded credentials; a database URL that does not parse is reported without its
text (it can hold the password). evalnever prints or stores the credentials of a--url(***@in errors, traces and
results).- Thread ids are one namespace: the server generates random ids when a client names none,
and the docs say to use unguessable ones (an id another principal used first is theirs).