You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
kagent → AX → Substrate: Making AX the Runtime Boundary for Agent Applications
#447
kagent → AX → Substrate: Making AX the Runtime Boundary for Agent Applications
I've implemented and deployed an integration that changes kagent's runtime dependency from kagent → Substrate to kagent → AX → Substrate.
The stack runs on Kubernetes with gVisor. In the demo, I create an Agent through the kagent UI, let AX prepare its runtime, and start a conversation with a DeepSeek-backed agent.
The main idea is to make AX the runtime interface that kagent depends on: kagent owns the application and its sessions; AX owns the runtime contract; Substrate provides the underlying execution mechanisms.
The additions described below are part of this integration, not features I am claiming are already available in upstream AX. The repository links point to development branches and may change after this post.
1. Why introduce AX between kagent and Substrate?
In the kagent version I started from, kagent directly used Substrate to create and run agent instances. Substrate already provided the essential mechanisms: Actors, WorkerPools, isolation, snapshots, and suspend/resume.
However, kagent also had to translate application requirements into those infrastructure resources. Creating a conversational agent involved constructing an ActorTemplate, preparing a golden snapshot, selecting a WorkerPool, creating an Actor, and handling backend routing. Checkpoint/fork and sandbox operations introduced further dependencies on snapshot semantics and the Guest protocol.
AX's existing Task abstraction offered a starting point for separating them. Rather than asking every agent application to implement its own Substrate integration, AX can expose a runtime contract that applications share.
That is the value I wanted to explore: an application should be able to ask AX to prepare, create, access, and restore an agent instance without constructing Substrate resources itself.
Adding AX also adds a service boundary and operational complexity. The abstraction needs to cover the whole interaction surface to justify that cost. Moving Actor creation behind AX while leaving WorkerPool management and backend routing in kagent would only partially solve the problem.
2. Goals and responsibilities
The integration has four goals:
Make AX models and APIs the runtime dependency exposed to kagent.
Cover runtime preparation, capacity, lifecycle, and access together.
Extend AX while retaining its ordinary Task execution path.
The intended division of responsibilities is:
Layer
What it owns
kagent
What the Agent does, who can use it, and how its sessions and application history work
AX
How the runtime is prepared, which capacity it uses, how an instance is identified and accessed, and how its lifecycle is managed
Substrate
How Actors execute, how worker capacity is provided, and how isolation, snapshots, and backend networking work
This does not make backend changes free. It gives those changes a defined integration boundary in AX instead of spreading them through kagent's business services.
3. Architecture
flowchart TB
subgraph K["kagent — application layer"]
Config["Harness + AgentTemplate<br/>Compile a runtime revision"]
Session["Session / Sandbox<br/>Lifecycle and checkpoint / fork"]
Dialer["RuntimeDialer<br/>A2A access"]
end
subgraph A["AX — runtime interface"]
Prepared["PreparedRuntime<br/>Prepare and reuse"]
Group["TaskGroup<br/>Shared capacity"]
Task["Task<br/>Runtime instance"]
Gateway["Shared TaskGateway<br/>Route by Task UID"]
Prepared -->|"Referenced by"| Task
Group -->|"Hosts"| Task
Gateway -->|"Targets"| Task
end
subgraph S["Substrate — execution backend"]
Template["ActorTemplate + golden snapshot"]
Pool["WorkerPool"]
Actor["Actor"]
end
Config -->|"Compile and prepare"| Prepared
Config -->|"Select group"| Group
Session -->|"AX lifecycle APIs"| Task
Dialer -->|"A2A over mTLS"| Gateway
Prepared --> Template
Group -->|"1:1"| Pool
Task -->|"1:1"| Actor
Loading
The diagram shows logical relationships, not every network hop. Backend resource mappings stay inside AX. kagent does not construct ActorTemplates, manage WorkerPools, or address Actors directly.
4. The key design choices
TaskGroup: replace direct WorkerPool management
TaskGroup represents shared runtime capacity. In this implementation, one TaskGroup maps to one Substrate WorkerPool.
Its configuration is deliberately small:
Field
Meaning
replicas
Number of workers, not the number of Tasks or Agents
sandboxClass
The platform's supported isolation class, such as gvisor
snapshotLocation
Snapshot storage location
Multiple Harnesses and sessions can share a TaskGroup. Each Task binds to its group using a name and UID, preventing an existing reference from silently attaching to a different group recreated under the same name.
I initially considered a batch-oriented resource, but the semantics did not fit. Sessions arrive over time. The group should provide capacity, not maintain a fixed number of Tasks or replace a deleted session.
The implementation includes scaling checks and reference protection for deletion. TaskGroups can be created while AX Server is running; they do not have to be declared at startup. The CLI currently provides ax group create|get|list|scale|delete. TaskGroup is an AX API resource, not a Kubernetes CRD, and ax apply does not yet handle it.
PreparedRuntime: prepare an environment before creating sessions
PreparedRuntime separates preparing an Agent's environment from creating a particular session instance.
kagent compiles a Harness and AgentTemplate into an immutable revision. It then calls AX to prepare the image, command, runtime configuration, readiness requirements, and credential bindings. AX translates that configuration into the backend template and golden snapshot—the prepared starting state used to instantiate runtimes.
The lifecycle is Prepare / Get / Release. kagent publishes Agent readiness only after preparation succeeds and the usable revision is persisted. Old revisions are released through reference-aware cleanup.
Many sessions can reuse a prepared starting environment, while each session retains its own runtime state and identity. Creating an Agent therefore does not immediately create a user session's Task. A golden Actor used during preparation is separate from a session instance.
Managed Tasks: replace direct Actor lifecycle operations
For the conversation path, one kagent Session binds to one AX Task, backed by one Actor. An A2A Task still represents application work within that session; it is not the same thing as the AX runtime instance.
The managed path adds groupRef, preparedRuntimeRef, optional restoreFrom, and separate runtime status. It covers create, resume, pause, suspend, delete, and checkpoint operations.
Ordinary Tasks retain their existing path. Persisted ownership distinguishes managed resources, so omitting managed fields from a later request cannot bypass the managed-resource checks.
Shared TaskGateway: hide backend addressing
kagent's RuntimeDialer sends A2A traffic to a shared TaskGateway inside AX Server. The gateway validates access, identifies the target Task by UID, resumes it when needed, and forwards requests through Substrate's routing path.
This is an added service capability within AX Server, not a separate gateway deployment for each group. A stateful conversation must reach its own Task rather than an arbitrary instance in the group.
The gateway supports A2A unary calls, streaming, and cancellation. Sandbox process and filesystem operations use AX's TaskExecutionService, with AX-owned messages and backend Guest protocol conversion inside AX.
Together, these changes move the dependency boundary beyond simply replacing an Actor reference with a Task reference.
5. What happens during a conversation?
Prepare the Agent
An operator creates or selects a TaskGroup. In kagent, a Harness chooses the adapter, runtime image, and group; an AgentTemplate defines the model, prompt, and tools; an Agent combines the two.
sequenceDiagram
actor U as User / operator
participant K as kagent
participant A as AX
participant S as Substrate
U->>A: Create or select TaskGroup
A->>S: Ensure corresponding worker capacity
U->>K: Configure model, Harness, template, and Agent
K->>A: Resolve TaskGroup UID
K->>K: Compile immutable runtime revision
K->>A: PrepareRuntime with stable request identity
A->>S: Prepare template and golden snapshot
loop Observe preparation
K->>A: GetPreparedRuntime
A-->>K: Preparation status
end
Note over K,A: Successful preparation
K->>K: Persist usable revision and AX references
K-->>U: Agent Ready
Loading
No manual PreparedRuntime creation is needed: the kagent controller drives this process.
Create the session and handle the first message
The current UI creates the Session when the user submits the first message. Internally, session creation and message dispatch are separate steps: kagent records the creation intent, creates the AX Task, and persists its binding before dispatching the message.
sequenceDiagram
actor U as User
participant K as kagent / TaskStore
participant A as AX / TaskGateway
participant S as Substrate
participant R as ADK runtime
participant M as Model API
U->>K: Start a conversation with a message
K->>K: Record Session and stable creation intent
K->>A: CreateTask with group and runtime references
A->>S: Create suspended Actor and runtime bindings
A-->>K: Task identity
K->>K: Persist Session-to-Task binding
K->>A: Dispatch A2A request to Task UID
A->>A: Validate caller and target
opt Task needs activation
A->>S: Resume instance
S->>R: Restore runtime and check readiness
end
A->>R: Route A2A request through backend
R->>K: HTTPS TaskStore call with runtime credential
K->>A: Authenticate runtime
A-->>K: Verified Task identity
K->>K: Validate Session binding
K-->>R: Session identity and requested business state
R->>R: Use instance-local ADK conversation state
R->>M: Call model through configured egress
M-->>R: Model output
R-->>A: A2A events
A-->>K: Forward events
K-->>U: Display response
R->>K: Persist business state and coordinate completion
K->>A: Request lifecycle transition when appropriate
A->>S: Perform backend transition
K->>A: Observe completion
A-->>K: Runtime state and boundary reference
K->>K: Persist boundary
Loading
This is a logical sequence. Streaming and state writes may interleave, and subsequent TaskStore requests still require authentication. kagent's public business history is distinct from the ADK runtime's private conversation state, stored in the instance's durable directory.
A later message in the same Session targets the same Task UID. If the Task is suspended, AX restores it; otherwise, it routes to the running instance. A boundaryRef identifies the recorded runtime boundary without making kagent handle backend snapshot addresses.
Checkpoint/fork builds on that separation: kagent owns the business history cutoff and creates a new Session, while AX creates a new Task from a validated restore reference. Deleting a Session cleans up its Task without deleting shared capacity or prepared runtimes that remain referenced.
6. The working demo
Environment
The demo runs on an existing Kubernetes cluster with:
Substrate using gVisor (sandboxClass: gvisor).
The extended AX Server and persistent Redis.
The paired kagent controller, UI, CRDs, and ADK runtime image, with PostgreSQL.
mTLS access to AX, HTTPS runtime callbacks, and the outbound credential and trust configuration needed for model calls.
The application path requires more than healthy Pods: model access, runtime credential injection, and TaskStore callbacks must also work.
With the AX endpoint and client certificates configured:
ax --server "$AX_ENDPOINT" --atespace kagent group create \
--file "$DEMO_DIR/task-group-demo.yaml" \
--request-id create-kagent-demo
ax --server "$AX_ENDPOINT" --atespace kagent group get \
--name kagent-task-group-demo
The snapshot location is specific to this deployment. Retrying the same creation uses the same request identity; a different operation should not reuse it.
The creation returned Provisioning, followed by Ready with one ready worker:
The corresponding WorkerPool and worker Pod were ready. This is an operator-side inspection of the mapping, not a direct dependency in kagent's application path:
Configure the Agent through the UI
The actual demo used these resources:
Resource
Demo configuration
ModelConfig
deepseek-v4-pro, using the OpenAI-compatible provider and model ID deepseek-v4-pro
Model endpoint
https://api.deepseek.com
Harness
kagent-harness-demo, with the kagent adapter and a digest-pinned ADK image
TaskGroup
kagent-task-group-demo
AgentTemplate
kagent-template-demo, with a general-knowledge assistant prompt
Agent
deepseek-agent-demo, referencing that template and Harness
The model API key is stored in a Kubernetes Secret. The runtime receives a credential binding rather than embedding the actual key in the shared runtime configuration.
The Harness form selects AX capacity directly. This particular demo also set a snapshot-location override; leaving it unset would use the group's location.
The template supplies the model and prompt:
The Agent references both resources:
After AX preparation completed, the Agent appeared as Ready:
Start a conversation and inspect its runtime
I opened the Agent and sent two messages: “who are you?” and “what can you do?” The UI displayed responses from the configured agent:
Inspecting the Session's AX Task showed its UID-bound TaskGroup and PreparedRuntime references, together with a recorded suspend boundary:
The Task output contains executionVersion: 3, an idle Suspend operation, a boundaryRef, and restorable: true. This is a post-execution observation; it should not be mistaken for a capture immediately after initial Task creation.
The corresponding PreparedRuntime query reported kind: Service and phase: Ready. The following selected fields show the runtime-to-group relationship:
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
kagent → AX → Substrate: Making AX the Runtime Boundary for Agent Applications
I've implemented and deployed an integration that changes kagent's runtime dependency from kagent → Substrate to kagent → AX → Substrate.
The stack runs on Kubernetes with gVisor. In the demo, I create an Agent through the kagent UI, let AX prepare its runtime, and start a conversation with a DeepSeek-backed agent.
The main idea is to make AX the runtime interface that kagent depends on: kagent owns the application and its sessions; AX owns the runtime contract; Substrate provides the underlying execution mechanisms.
Implementation repositories:
The additions described below are part of this integration, not features I am claiming are already available in upstream AX. The repository links point to development branches and may change after this post.
1. Why introduce AX between kagent and Substrate?
In the kagent version I started from, kagent directly used Substrate to create and run agent instances. Substrate already provided the essential mechanisms: Actors, WorkerPools, isolation, snapshots, and suspend/resume.
However, kagent also had to translate application requirements into those infrastructure resources. Creating a conversational agent involved constructing an ActorTemplate, preparing a golden snapshot, selecting a WorkerPool, creating an Actor, and handling backend routing. Checkpoint/fork and sandbox operations introduced further dependencies on snapshot semantics and the Guest protocol.
This left two kinds of logic in kagent:
AX's existing Task abstraction offered a starting point for separating them. Rather than asking every agent application to implement its own Substrate integration, AX can expose a runtime contract that applications share.
That is the value I wanted to explore: an application should be able to ask AX to prepare, create, access, and restore an agent instance without constructing Substrate resources itself.
Adding AX also adds a service boundary and operational complexity. The abstraction needs to cover the whole interaction surface to justify that cost. Moving Actor creation behind AX while leaving WorkerPool management and backend routing in kagent would only partially solve the problem.
2. Goals and responsibilities
The integration has four goals:
The intended division of responsibilities is:
This does not make backend changes free. It gives those changes a defined integration boundary in AX instead of spreading them through kagent's business services.
3. Architecture
flowchart TB subgraph K["kagent — application layer"] Config["Harness + AgentTemplate<br/>Compile a runtime revision"] Session["Session / Sandbox<br/>Lifecycle and checkpoint / fork"] Dialer["RuntimeDialer<br/>A2A access"] end subgraph A["AX — runtime interface"] Prepared["PreparedRuntime<br/>Prepare and reuse"] Group["TaskGroup<br/>Shared capacity"] Task["Task<br/>Runtime instance"] Gateway["Shared TaskGateway<br/>Route by Task UID"] Prepared -->|"Referenced by"| Task Group -->|"Hosts"| Task Gateway -->|"Targets"| Task end subgraph S["Substrate — execution backend"] Template["ActorTemplate + golden snapshot"] Pool["WorkerPool"] Actor["Actor"] end Config -->|"Compile and prepare"| Prepared Config -->|"Select group"| Group Session -->|"AX lifecycle APIs"| Task Dialer -->|"A2A over mTLS"| Gateway Prepared --> Template Group -->|"1:1"| Pool Task -->|"1:1"| ActorThe diagram shows logical relationships, not every network hop. Backend resource mappings stay inside AX. kagent does not construct ActorTemplates, manage WorkerPools, or address Actors directly.
4. The key design choices
TaskGroup: replace direct WorkerPool management
TaskGroup represents shared runtime capacity. In this implementation, one TaskGroup maps to one Substrate WorkerPool.
Its configuration is deliberately small:
replicassandboxClassgvisorsnapshotLocationMultiple Harnesses and sessions can share a TaskGroup. Each Task binds to its group using a name and UID, preventing an existing reference from silently attaching to a different group recreated under the same name.
I initially considered a batch-oriented resource, but the semantics did not fit. Sessions arrive over time. The group should provide capacity, not maintain a fixed number of Tasks or replace a deleted session.
The implementation includes scaling checks and reference protection for deletion. TaskGroups can be created while AX Server is running; they do not have to be declared at startup. The CLI currently provides
ax group create|get|list|scale|delete. TaskGroup is an AX API resource, not a Kubernetes CRD, andax applydoes not yet handle it.PreparedRuntime: prepare an environment before creating sessions
PreparedRuntime separates preparing an Agent's environment from creating a particular session instance.
kagent compiles a Harness and AgentTemplate into an immutable revision. It then calls AX to prepare the image, command, runtime configuration, readiness requirements, and credential bindings. AX translates that configuration into the backend template and golden snapshot—the prepared starting state used to instantiate runtimes.
The lifecycle is Prepare / Get / Release. kagent publishes Agent readiness only after preparation succeeds and the usable revision is persisted. Old revisions are released through reference-aware cleanup.
Many sessions can reuse a prepared starting environment, while each session retains its own runtime state and identity. Creating an Agent therefore does not immediately create a user session's Task. A golden Actor used during preparation is separate from a session instance.
Managed Tasks: replace direct Actor lifecycle operations
For the conversation path, one kagent Session binds to one AX Task, backed by one Actor. An A2A Task still represents application work within that session; it is not the same thing as the AX runtime instance.
The managed path adds
groupRef,preparedRuntimeRef, optionalrestoreFrom, and separate runtime status. It covers create, resume, pause, suspend, delete, and checkpoint operations.Ordinary Tasks retain their existing path. Persisted ownership distinguishes managed resources, so omitting managed fields from a later request cannot bypass the managed-resource checks.
Shared TaskGateway: hide backend addressing
kagent's RuntimeDialer sends A2A traffic to a shared TaskGateway inside AX Server. The gateway validates access, identifies the target Task by UID, resumes it when needed, and forwards requests through Substrate's routing path.
This is an added service capability within AX Server, not a separate gateway deployment for each group. A stateful conversation must reach its own Task rather than an arbitrary instance in the group.
The gateway supports A2A unary calls, streaming, and cancellation. Sandbox process and filesystem operations use AX's TaskExecutionService, with AX-owned messages and backend Guest protocol conversion inside AX.
Together, these changes move the dependency boundary beyond simply replacing an Actor reference with a Task reference.
5. What happens during a conversation?
Prepare the Agent
An operator creates or selects a TaskGroup. In kagent, a Harness chooses the adapter, runtime image, and group; an AgentTemplate defines the model, prompt, and tools; an Agent combines the two.
sequenceDiagram actor U as User / operator participant K as kagent participant A as AX participant S as Substrate U->>A: Create or select TaskGroup A->>S: Ensure corresponding worker capacity U->>K: Configure model, Harness, template, and Agent K->>A: Resolve TaskGroup UID K->>K: Compile immutable runtime revision K->>A: PrepareRuntime with stable request identity A->>S: Prepare template and golden snapshot loop Observe preparation K->>A: GetPreparedRuntime A-->>K: Preparation status end Note over K,A: Successful preparation K->>K: Persist usable revision and AX references K-->>U: Agent ReadyNo manual PreparedRuntime creation is needed: the kagent controller drives this process.
Create the session and handle the first message
The current UI creates the Session when the user submits the first message. Internally, session creation and message dispatch are separate steps: kagent records the creation intent, creates the AX Task, and persists its binding before dispatching the message.
sequenceDiagram actor U as User participant K as kagent / TaskStore participant A as AX / TaskGateway participant S as Substrate participant R as ADK runtime participant M as Model API U->>K: Start a conversation with a message K->>K: Record Session and stable creation intent K->>A: CreateTask with group and runtime references A->>S: Create suspended Actor and runtime bindings A-->>K: Task identity K->>K: Persist Session-to-Task binding K->>A: Dispatch A2A request to Task UID A->>A: Validate caller and target opt Task needs activation A->>S: Resume instance S->>R: Restore runtime and check readiness end A->>R: Route A2A request through backend R->>K: HTTPS TaskStore call with runtime credential K->>A: Authenticate runtime A-->>K: Verified Task identity K->>K: Validate Session binding K-->>R: Session identity and requested business state R->>R: Use instance-local ADK conversation state R->>M: Call model through configured egress M-->>R: Model output R-->>A: A2A events A-->>K: Forward events K-->>U: Display response R->>K: Persist business state and coordinate completion K->>A: Request lifecycle transition when appropriate A->>S: Perform backend transition K->>A: Observe completion A-->>K: Runtime state and boundary reference K->>K: Persist boundaryThis is a logical sequence. Streaming and state writes may interleave, and subsequent TaskStore requests still require authentication. kagent's public business history is distinct from the ADK runtime's private conversation state, stored in the instance's durable directory.
A later message in the same Session targets the same Task UID. If the Task is suspended, AX restores it; otherwise, it routes to the running instance. A
boundaryRefidentifies the recorded runtime boundary without making kagent handle backend snapshot addresses.Checkpoint/fork builds on that separation: kagent owns the business history cutoff and creates a new Session, while AX creates a new Task from a validated restore reference. Deleting a Session cleans up its Task without deleting shared capacity or prepared runtimes that remain referenced.
6. The working demo
Environment
The demo runs on an existing Kubernetes cluster with:
sandboxClass: gvisor).The application path requires more than healthy Pods: model access, runtime credential injection, and TaskStore callbacks must also work.
Create a TaskGroup
The group used in the demo was:
With the AX endpoint and client certificates configured:
The snapshot location is specific to this deployment. Retrying the same creation uses the same request identity; a different operation should not reuse it.
The creation returned Provisioning, followed by Ready with one ready worker:
The corresponding WorkerPool and worker Pod were ready. This is an operator-side inspection of the mapping, not a direct dependency in kagent's application path:
Configure the Agent through the UI
The actual demo used these resources:
deepseek-v4-pro, using the OpenAI-compatible provider and model IDdeepseek-v4-prohttps://api.deepseek.comkagent-harness-demo, with thekagentadapter and a digest-pinned ADK imagekagent-task-group-demokagent-template-demo, with a general-knowledge assistant promptdeepseek-agent-demo, referencing that template and HarnessThe model API key is stored in a Kubernetes Secret. The runtime receives a credential binding rather than embedding the actual key in the shared runtime configuration.
The Harness form selects AX capacity directly. This particular demo also set a snapshot-location override; leaving it unset would use the group's location.
The template supplies the model and prompt:
The Agent references both resources:
After AX preparation completed, the Agent appeared as Ready:
Start a conversation and inspect its runtime
I opened the Agent and sent two messages: “who are you?” and “what can you do?” The UI displayed responses from the configured agent:
Inspecting the Session's AX Task showed its UID-bound TaskGroup and PreparedRuntime references, together with a recorded suspend boundary:
The Task output contains
executionVersion: 3, an idle Suspend operation, aboundaryRef, andrestorable: true. This is a post-execution observation; it should not be mistaken for a capture immediately after initial Task creation.The corresponding PreparedRuntime query reported
kind: Serviceandphase: Ready. The following selected fields show the runtime-to-group relationship:{ "metadata": { "name": "deepseek-agent-demo-5fb960b71a51", "atespace": "kagent", "uid": "24bfc8f1265b8ce5068360f39ae20828" }, "spec": { "groupRef": { "atespace": "kagent", "name": "kagent-task-group-demo", "uid": "5cde9957076ba25dfb2cb18bcc0477ed" }, "kind": "Service" }, "phase": "Ready" }All reactions