-
Notifications
You must be signed in to change notification settings - Fork 0
Steps
There is no agents: block. A named agent profile turned out to be doing
five separate jobs, which is why it could not simply be renamed. Each has a
home now, and none of them is an agent:
| the old job | where it lives |
|---|---|
which model to run (provider/model) |
a fleet — model: on the step (Model-Selection) |
execution cost cap (budget) |
the runtime (Cost-Accounting) |
| private memory namespace | an opaque scope key, defaulting to the step identity (Memory) |
live-session pool (session) |
(runtime, model, key) (Runtimes, below) |
| self-improvement track record | an opaque track-record key, defaulting to the step identity (Outcomes) |
| behavior (guidance, skill, workspace, timeouts, archive, host, isolation) | the step, shared with a YAML anchor |
Existing configs are migrated automatically at boot (or by
conductor config migrate), and the migration is careful to preserve your
accumulated history — see Migration below.
Behavior lives on the step that dispatches the work. To share it between
steps, park a base under any top-level x- key — the loader ignores
x--prefixed sections (the docker-compose extension-field convention), so
they exist purely to hold anchors — and pull it in one of two ways:
-
<<: *base— plain YAML merge. Dumb: the step's own key wins outright, scalars and lists alike. Use it for a copy of a base. -
extends: *base— conductor's own merge. Field-aware: scalars override, lists append, maps deep-merge, andguidance:stacks. Use it to add to a base.!overrideand!resetare the escape hatches.
Both work identically in a Packs. See Reuse for the
full rules; the example below uses <<:.
x-templates:
fixer: &fixer
type: agent
model: claude-opus-5 # a fleet name, a model id, a wildcard,
# or { any: [...], required: bool }
thinking: "" # runtime launch hint (optional)
mode: "" # runtime session mode (optional)
workspace: worktree # local | worktree
# workspace: { isolation: local, pin: triage } # …or reuse a NAMED workspace
wait_timeout: 30m
archive_when_done: true # archived inline the moment the step finishes
labels: { team: autopilot }
# runtime: paseo # pin the backend (default: the default: true runtime)
# host: build-box # a [[Hosts]] SSH target its runtime launches on
# guidance: | # tone/format; STACKS on the scoped baseline
# One or two sentences, plain and direct.
# memory: true # shared-memory injection ([[Memory]])
# outcome_feedback: true # append this step's track record ([[Outcomes]])
# outcome_key: reviewers # …or pool several steps onto one record
# session: { key: "{{.repo}}#{{.pr}}" } # this step's own session pool
# skill: { verbs: [gh.comment] } # reach back into conductor ([[Agent-Skill]])
# isolation: { mode: namespace } # per-dispatch sandboxing ([[Isolation]])
triggers:
github.pull_request:
steps:
- <<: *fixer # merge the anchor…
id: fix # …then this step's own fields, which win
prompt: "Resolve the conflict on {{.repo}}#{{.pr}}."A step may of course carry these fields inline — an anchor is only how you avoid retyping them.
Two things to know:
-
Anchors are file-local. YAML resolves them per document, so an anchor
defined in
config.yamlis not visible in an importedconf.d/*.yaml. For reuse across files, useextends:on a map section (runtimes:/workflows:/handoffs:) or a named step, below. See Reuse. -
An anchor copies fields, not identity. Two steps merging one anchor
are still two identities.
name:is the separate, deliberate opt-in to sharing one — see below.
There is no top-level steps: section. A step lives in the workflow or
trigger that runs it, and anything that must point at a particular one
addresses it there:
review-flow/architect by the step's own id: (or name:, if it has no id:)
review-flow[2] by position, for a step that carries neither
The container is a workflows: entry or a named trigger. That is the
same slot the identity ladder uses below — id if present, else index — so
a reference and an identity always name the same step by the same rule:
review-flow/architect addresses exactly the step whose structural
identity is workflow:review-flow/architect.
Two things use references:
- a [[Teams|
team:]] block, whoseplanner/worker/critic/reconcileroles each name a step. The role takes that step's behavior and its identity, so every team pointing at it shares one memory namespace, session pool, and track record; - a Packs overlay, where a consumer overrides one step of a pack
by reference (
packs.review.steps: { review-flow/review: { … } }).
workflows:
roles:
steps:
- id: architect
type: agent
guidance: "Decompose before building. Name the risky part."
prompt: "…"
triggers:
- on: gh.issue_matched
steps:
- id: feature
prompt: "Implement {{.url}}"
team: { planner: roles/architect, worker: roles/architect }The index form is the escape hatch, not the habit: it is positional, so
inserting a step above shifts it. Give a step an id: and reference that.
agent: still parses, but it is now a free-form attribution label that
selects nothing. What identifies a dispatch is its identity, below.
Memory scoping, session affinity, and outcome tracking all default their key to the step's identity. It must be stable across restarts, so it is a pure function of config — never a per-run value:
- an explicit
name:— author-pinned, and shareable: two steps with the same name share one memory namespace, one session pool, and one track record — exactly the reuse a sharedagent: fixergave you. Ateam:role inherits the name of the step it references, for the same reason. Merging an anchor does not do this: an anchor copies fields and leaves identity alone. -
structural — the enclosing qualified trigger/workflow plus the step's
slot: its
id:, else itsname:, else its index in the list —github.pull_request/security, orgithub.pull_request/2. This is the default, and it is the same slot a step reference uses. - a deterministic fingerprint of the step's definition, for a step with no enclosing context.
Editing a step's prompt does not rotate its identity. Reordering may, for
a step that has neither a name: nor an id: — give it one to pin it.
id:is not rung 1. It is the run-local handlesteps.<id>.outputs.*addresses and it is near-universal, so treating it as a global identity would make two unrelated triggers that both useid: fixsilently share a track record. Insteadid:supplies the structural slot, which is what makes structural identity survive reordering.
A step sets exactly one form — what it does:
| Form | Spelling |
|---|---|
| dispatch an agent |
type: agent (or a bare prompt: with agent:) |
| ask typed questions |
decide: + state: / questions: — answered with probabilities by a decision runtime or any agent runtime; see Decide-Steps
|
| call a connector verb |
uses: <connector>.<verb> + options:
|
| run code |
use: <engine> + code:, or use: cli + command: — see Code-Steps
|
| run a program |
type: command + command:
|
| call a workflow |
call: <workflow> + with: — see Workflows
|
| a helper step |
sleep:, log:, set:, assert:, fail:, wait_for: — conductor's own flow control, below |
Two of those spellings moved, and both old ones still parse:
-
use:selects the code ENGINE. It resolves like a connector's or a runtime'suse:— the builtin (cli), then the official plugin repo'sengines/<name>, which is where the scripting engines (js,go-embed,risor,lua) now live, then an explicit repo or path — and it also takes a host interpreter by name (bash,python3) or by path.run:is the same key and everyrun: js/run: bashconfig keeps working; set one or the other, not both. -
call:is the workflow call. A step-leveluse:used to mean this.conductor config migrate(and the boot auto-migration) rewrites everyuse: <workflow>tocall: <workflow>;workflow:remains valid as the older spelling of the same field.
steps:
- { id: triage, use: js, code: "return { sev: ctx.body.severity }" }
- { id: build, use: cli, command: [make, -C, ./svc, release] }
- { id: review, call: review-flow, with: { pr: "{{.pr}}" } }A helper step is a form conductor runs itself: no agent, no engine, no
connector, no command behind it. There is nothing to dispatch, nothing to
install and nothing to grant — which is the point, because the alternative
was a use: cli step shelling out to a coreutil just to pause a flow.
The family, each a form in its own right:
| Helper | Does |
|---|---|
sleep: <duration> |
Pause the flow for a positive duration. |
log: <message> |
Render a message (templated) into the run log — a breadcrumb. |
set: {<key>: <value>} |
Publish computed values as this step's outputs, read downstream as {{.<id>.<key>}}. |
assert: <expr> |
Fail the run unless the condition is truthy — the same grammar as if:. |
fail: <message> |
Stop the run with a rendered message; guard it with if: for a conditional abort. |
wait_for: {…} |
Poll a read verb until a condition holds, or a timeout — below. |
Each is mutually exclusive with type: / use: / uses: / call: and with
the other helpers — a step does one thing — and each takes id:, if:,
for_each: and step hooks like any other step, appears in the run timeline,
and runs the same under [[One-Shot|conductor once]] as on the daemon.
steps:
- log: "addressing {{.repo}}#{{.pr}}" # a breadcrumb in the run log
- { id: bail, if: "!run_id", fail: "no run_id on this event" }
- assert: "checks_passed && !draft" # stop unless the guard holds
- id: vars
set: { branch: "fix/{{.pr}}", attempts: 1 } # {{.vars.branch}} downstream
- sleep: 5s # pause, cancellation-awareWhat to know:
-
sleep's duration is the one every other field takes —500ms,5s,30m,7d,1d12h(Configuration) — and must be positive;sleep: 0and a negative are refused at load. Cancellation cuts it short: a shutdown, a steptimeout:, or a budget cut-off leaves asleep: 30mat once. A dry run prints[dry-run] would sleep 5sand moves on. -
setvalues keep their type.attempts: 1reads back as the number1, not"1", and a value lifted from an earlier step's list/map arrives intact. Give the step anid:— that is how the values are addressed. -
assertandfailboth stop the run on failure, the same way any step error does (the run reportsworkflow_failed).assertis the conditional form — it fails unless the expr is truthy — andfailis unconditional, so it usually sits behind anif:. Both render/evaluate in a dry run too, the wayif:does. -
In an agent-authored plan each helper is a
class an agent may only emit if the operator listed it in
policy.agent_authored.verbs. They reach nothing outside conductor, but they still spend the run — admitting one stays a decision.
Polls a read verb until its output satisfies a condition, or a timeout
elapses. It is the principled form of a bare sleep: before a step that
depends on state settling — instead of guessing how long a cancellation takes
to land, wait for the run to actually reach a terminal state:
steps:
- uses: gh.cancel_run
options: { repo: "{{.repo}}", run_id: "{{.run_id}}" }
- wait_for:
uses: gh.get_run # the read verb to poll
options: { repo: "{{.repo}}", run_id: "{{.run_id}}" }
until: "status == 'completed'" # condition over its outputs
every: 10s # poll interval (default 10s)
timeout: 2m # give up after this — required
id: settle
- uses: gh.rerun_run
options: { repo: "{{.repo}}", run_id: "{{.run_id}}" }-
until:reads the polled verb's outputs at the top level — agh.get_runreturnsstatus,conclusion, … sountil: "status == 'completed'"sees them directly, alongside every fact ({{.repo}}) and prior-step output. Same expression grammar asif:. -
The verb's
options:are re-templated each poll, exactly as auses:step would render them. -
On success the last poll's outputs become the step's outputs, so a
later step reads
{{.settle.conclusion}}. On timeout the step fails — soften a flow that should continue anyway withcontinue_on_error: true. -
timeout:is required and positive;every:defaults to10s. The whole wait is cancellation-aware — a shutdown drops out immediately, and the timeout is just another deadline on the same context. A dry run does one stubbed poll, says[dry-run] would poll …, and does not wait.
The family exists so the next helper is a small, predictable addition rather than a new subsystem. Anything that needs no agent, engine, verb or command to execute belongs here.
| Field | Meaning |
|---|---|
name |
Pins the step's identity (see above). Shareable on purpose. |
id |
The step's slot: how outputs are addressed (steps.<id>.outputs.*), what a step reference points at, and — absent a name: — the structural half of its identity. |
use |
The code ENGINE this step runs on (run: is the same key). See Code-Steps. |
command |
The argv for use: cli and for type: command — a list of words, or one string split on whitespace (quotes honored, no shell). |
call |
A workflow to run as this step, with with: for its inputs. See Workflows. |
sleep |
Pause the flow for a positive duration — a helper step, above. |
log |
Render a message into the run log — a helper step. |
set |
Publish computed values as the step's outputs ({{.<id>.<key>}}) — a helper step. |
assert |
Fail the run unless the expr is truthy — a helper step. |
fail |
Stop the run with a rendered message — a helper step. |
wait_for |
Poll a read verb (uses:) until until: holds or timeout: — a helper step. |
model |
Which model to run: a fleet name, a model id, a wildcard, an inline list, or { any, required }. Unset → the runtime's models.default:, then a bare launch. See Model-Selection. |
runtime |
A runtimes.<name> entry to run on (default: the default: true runtime, else the built-in paseo). See Runtimes. |
thinking / mode
|
Runtime launch hints, passed through where the runtime supports them. |
workspace |
local or worktree — the existing checkout, or a fresh git worktree. |
expect_push |
Mark a fixer step that must LAND a change on the target. A clean run that leaves work unlanded (a non-empty proposed diff never pushed) becomes a no_progress failure instead of a silent success — see Workflows hooks. Leave off for a review/judge step. |
wait_timeout |
How long a foreground dispatch waits before giving up. |
archive_when_done |
Whether the agent (and its conductor-created workspace) is archived the moment its step finishes. Forced off for a background: hand-off step. |
watch |
Reactive live step: run steps: every every: (a mini-workflow — fact reads + if-guarded actions). Actions: step.bail (reason gone → tear down), step.rerun (re-run this step on the new state), or a workflow: step (re-run a workflow — re-review). Works on ANY live step, not just hand-offs — e.g. a long-processing step can bail when its input disappears or rerun when new work arrives. (handoff.bail/handoff.rerun are accepted as deprecated aliases.) Operator-owned. See Hand-offs. |
idle_timeout |
Release an interactive hand-off still open after this long (drop the hold, archive the agent + workspace). The backstop for one nobody closed; the agent calling step.done is the precise signal. Off unless set. See Hand-offs. |
labels |
Extra key=value labels on the dispatched agent. |
host |
A Hosts SSH target this step's runtime launches on, overriding the runtime's own. |
guidance |
Tone/format that stacks on the scoped baseline ([[Policy|policy.guidance]]) rather than replacing it. A string, a list, or { replace: … }. See Reuse. |
memory |
Opt into shared-memory injection (Memory). true uses the run's context keys; a map { scopes, tags, limit } names arbitrary keys. Also gates writing: only an opted-in step may harvest its output into shared memory. |
session |
This step's own session pool: { key, idle_ttl, max_lifetime, end_on }, namespaced to its identity. A step with none joins the runtime's overall pool. See below. |
skill |
Verb tools + the secret broker over the daemon socket: { verbs, secrets_via, allow_secrets, max_calls }. See Agent-Skill. |
isolation |
Per-dispatch sandboxing. Wins over the runtime's own; needs a runtime conductor launches itself (not paseo). See Isolation. |
outcome_feedback |
true appends this step's own track record to its guidance. See Outcomes. |
outcome_key |
Override the track-record key (default: the identity), so several steps can pool one record. |
Note there is no budget here: a budget caps execution cost on a backend, so
it lives on the runtime (Cost-Accounting).
- Every step of a multi-step Workflows can differ — a common
pattern is a cheap
plannerstep triaging an issue and handing off to a strongerfixeronly when the triage justifies it (if: "{{.evaluate.has_context}} == true"). -
workspacegoverns the checkout lifecycle (persistent checkout vs. a fresh worktree per dispatch); the step's owncheckout:(checkout-pr|branch-off|none) governs what git state that checkout is put into. Independent knobs — aworkspace: worktreestep can still runcheckout: nonefor triage. -
workspacealso takes an object,{ isolation: local|worktree, pin: <name> }.pin:names a runtime workspace the step ALWAYS runs in — created on first use, reused by every later run — so a long-lived triage or chat agent finds its working directory as it left it. Without a pin, acheckout: nonestep gets a fresh workspace per dispatch that is archived when the run finishes. A pin cannot be combined with an explicitcheckout: checkout-prorbranch-off(conductor validaterejects it): those give each dispatch its own worktree, which is the opposite of reusing one. Withcheckout:unset the strategy comes from the trigger, so the pin applies to the runs with no repo context and is ignored on the ones that get a worktree. -
archive_when_done: truesteps are still protected from premature cleanup: reclaim skips one paused on a permission prompt, and an agent can hold itself alive with a.paseo-holdmarker in its worktree. - Which models a runtime can actually run is DISCOVERED, per runtime — see
Model-Discovery.
conductor validatereports arequired:fleet nothing can satisfy.
By default each dispatch gets a fresh agent. A session: block binds a live
session to the rendered key, so later events reach the same agent as a
follow-up with full prior context.
A binding is (runtime, resolvedModel, key). Both extra dimensions are
structural, not preferences: a live agent is one model on one runtime, and
you can resume neither a paseo session on codex nor an opus step into a haiku
session. A useful consequence — because a pack assigns a fleet per step, its
model assignments partition affinity for free.
It is declarable at two scopes:
runtimes:
paseo:
session: # the OVERALL pool: one agent per key,
key: "{{.repo}}#{{.pr}}" # shared by every step without its own
idle_ttl: 12h
end_on: [ gh._closed ]
x-templates:
reviewer: &reviewer
type: agent
session: # this step's OWN pool, namespaced to its
key: "{{.repo}}#{{.pr}}" # identity — so the same key string is
end_on: [ gh._closed ] # still a distinct sessionResolution per dispatch: the step's session: wins, else the
runtime's, else a fresh agent.
-
Serialized per key. At most one prompt in flight per session;
concurrent same-key events queue — the
group:one-run-per-key guarantee, extended across the session's life. -
Durable. The binding persists in conductor's own state
(
affinity.json, beside audit/dedup); after a restart the session resumes via the runtime's native handle (paseo re-binds the agent id, ACPsession/load). While bound, the agent is never reclaimed. -
Eviction. Idle past
idle_ttl, older thanmax_lifetime, or anend_onevent (matched as<connector>.<kind>, a bare<kind>, or<type>.<kind>, rendered against the event's own context). github's PR close and merge both arrive as the internal_closedsignal, soend_on: [ gh._closed ]is the evict-when-done form. -
_closedis eviction-only. It matchesend_on:but is not a trigger event —on: gh._closedfails validation. -
Runtime support. Needs a session-persistent runtime — paseo or ACP.
One-shot
cliruntimes fall back to a fresh agent per event; pair withmemory:(Memory) for continuity there. -
vs
group:and memory.group:batches a burst into one run; memory injects durable facts; session affinity reuses a live conversation. They compose.
Follow-up turns return Queued run refs: on paseo the prompt is delivered
via paseo send (no captured output for later steps); an ACP follow-up
returns the turn's output. A follow-up to a dead session is detected,
unbound, and replaced by a fresh spawn.
Bindings written before the re-key carry no runtime/model and are ignored on restore, so the next event starts a fresh session. Sessions are short-lived by design (24h idle default), so nothing durable is lost — a track record would have been another matter, which is why that one is preserved explicitly.
conductor config migrate (and the automatic boot migration) decomposes each
profile, and it is careful about one thing above all: your accumulated
history carries over. Memory, sessions, and outcomes used to key off the
agent NAME; they now key off the step IDENTITY. So the migration INLINES
each profile's behavior onto every step that referenced it, carrying
name: <the old agent name> — the identity ladder's top rung — onto
each one.
Where a profile was referenced from several steps IN THE SAME FILE, those
steps share a YAML anchor parked under x-migrated: instead of each
getting a copy. Where the steps are in different files, each gets its own:
anchors do not cross imports:, and duplicated config that works beats DRY
config that does not parse. Because agents: commonly sat in config.yaml
while the triggers naming it sat in conf.d/*.yaml, the profile table is
gathered from the whole import tree before any file is rewritten.
A profile nothing referenced is dropped with a note — there is no top-level
section left to park it in, and its behavior is in the .pre-connectors
backup.
The keys therefore come out identical to what your box already has on disk:
outcomeStats["fixer"] stays outcomeStats["fixer"], engagements still
match, and memories written under the old agent:fixer scope are still
recalled through a compatibility alias.
| Old | New |
|---|---|
agents.<n> |
its fields, inlined on each referencing step, with name: <n>
|
provider + model
|
model: (an exact pin — a migration never invents a fleet) |
provider alone |
nothing — that named a backend, not a model, so it becomes a bare launch |
budget |
runtimes.<the runtime it ran on>.budget |
controller |
runtime |
| behavior fields | the same key on each referencing step |
agents.<child>.extends: <parent> |
flattened — the parent's fields are copied into the child before inlining |
agent: <n> on a step |
the profile's fields, plus name: <n> (or <<: *n when shared in-file) |
Anything the decomposition has no home for is dropped with a note in the migration summary, never silently.
An agent step's final output may carry a ```plan block — steps in the
normal grammar that conductor validates, guards (policy.agent_authored,
Policy), and runs deterministically. Failures route back to this
agent's session (above) for a bounded revise loop, so a planning agent
should keep sessions; pairing with memory: lets it recall what worked.
The full loop — plan, choose from the catalog, supervise, promote — is in
Workflows.
A step and a runtime answer different questions. The step answers "what should run" — which
model, what tone, what workspace lifecycle. The runtime answers "how is it run" — which
process or API actually executes it. There is no provider:: a backend was never a property of
the work, so a step with model: claude-opus-5 can run on the built-in paseo dispatcher, on
agent-deck, or through opencode's HTTP API, unchanged, just by pointing runtime: at a
different entry — the model still routes through whichever runtime is selected. The one place
this decouples is an ACP or cli runtime: there, the runtime's own agent:/command:
names the tool directly (e.g. gemini over ACP), so the step's model: has nothing to route and
is ignored. With no runtimes: configured at all, every agent step runs on paseo, so this
distinction is invisible until you actually introduce a second runtime. See Runtimes for the
full resolution order and runtime kinds (Controllers is the legacy name).
Every agent conductor launches carries the built-in step connector's verbs
(the handoff.* spellings remain as deprecated aliases):
-
step.done— the agent's own completion signal, auto-granted to every dispatch. Calling it releases the agent and its conductor-created workspace (ledger-gated: an agent can only ever release itself — the target comes from its token, never an argument). On a step with anoutput_schema, the call ALSO delivers the result:conductor call step.done --output '<result>'is one atomic final action, validated against the schema at the verb boundary (an invalid value gets a precise, retryable error). Any JSON value shape the schema admits works — object, array, string, number, boolean, enum. -
step.bail/step.rerun—watch:rule actions on the LIVE step (not agent-callable): tear it down when its reason is gone, or supersede it with a re-run on the current state.
How agents learn about it: schema steps get the delivery instruction as their
schema directive; other foreground steps get a short "run step.done when
finished" footer; hand-offs get theirs in the hand-off guidance. Where the
signal is REACHABLE: paseo agents call the conductor CLI (creds injected in
env; remote hosts via the SSH tunnel); ACP and opencode runtimes get the same
surface as native MCP tools (step_done). A bare cli one-shot runtime has no
skill surface — there the step boundary itself is the completion signal
(conductor is blocked on the process), schema output rides the reply, and the
footer tells the agent to just finish normally if the command is unavailable.
Reclaim never depends on the call alone: a foreground step is archived at its
step boundary regardless, and hand-offs have idle_timeout as the backstop.
Setup
The model
- Connectors
- Workflows
- Reuse
- Settings-and-Templating
- Packs
- Verbs
- Code-Steps
- Stores
- Runtimes
- Model-Selection
- Model-Discovery
- Steps
- Decide-Steps
- Grouping
- Memory
- Binary-Data
- Agent-Skill
- Policy
- Gates
- Teams
- Outcomes
- Cost-Accounting
- Secrets
- Hosts
- Isolation
- Trust-and-Isolation
Connectors
Operations
- One-Shot
- Callable-Service
- Runs
- Hand-offs
- Notifications
- Migration
- Controllers (legacy name → Runtimes)