Skip to content
github-actions[bot] edited this page Sep 25, 2026 · 11 revisions

Steps (the retired agents: profile)

There is no agents: block. A named agent profile turned out to be doing five separate jobs, which is why it could not simply be renamed. Each has a home now, and none of them is an agent:

the old job where it lives
which model to run (provider/model) a fleet — model: on the step (Model-Selection)
execution cost cap (budget) the runtime (Cost-Accounting)
private memory namespace an opaque scope key, defaulting to the step identity (Memory)
live-session pool (session) (runtime, model, key) (Runtimes, below)
self-improvement track record an opaque track-record key, defaulting to the step identity (Outcomes)
behavior (guidance, skill, workspace, timeouts, archive, host, isolation) the step, shared with a YAML anchor

Existing configs are migrated automatically at boot (or by conductor config migrate), and the migration is careful to preserve your accumulated history — see Migration below.

Sharing step behavior

Behavior lives on the step that dispatches the work. To share it between steps, park a base under any top-level x- key — the loader ignores x--prefixed sections (the docker-compose extension-field convention), so they exist purely to hold anchors — and pull it in one of two ways:

  • <<: *base — plain YAML merge. Dumb: the step's own key wins outright, scalars and lists alike. Use it for a copy of a base.
  • extends: *base — conductor's own merge. Field-aware: scalars override, lists append, maps deep-merge, and guidance: stacks. Use it to add to a base. !override and !reset are the escape hatches.

Both work identically in a Packs. See Reuse for the full rules; the example below uses <<:.

x-templates:
  fixer: &fixer
    type: agent
    model: claude-opus-5              # a fleet name, a model id, a wildcard,
                                      #   or { any: [...], required: bool }
    thinking: ""                      # runtime launch hint (optional)
    mode: ""                          # runtime session mode (optional)
    workspace: worktree               # local | worktree
    # workspace: { isolation: local, pin: triage }   # …or reuse a NAMED workspace
    wait_timeout: 30m
    archive_when_done: true           # archived inline the moment the step finishes
    labels: { team: autopilot }
    # runtime: paseo                  # pin the backend (default: the default: true runtime)
    # host: build-box                 # a [[Hosts]] SSH target its runtime launches on
    # guidance: |                     # tone/format; STACKS on the scoped baseline
    #   One or two sentences, plain and direct.
    # memory: true                    # shared-memory injection ([[Memory]])
    # outcome_feedback: true          # append this step's track record ([[Outcomes]])
    # outcome_key: reviewers          #   …or pool several steps onto one record
    # session: { key: "{{.repo}}#{{.pr}}" }   # this step's own session pool
    # skill: { verbs: [gh.comment] }  # reach back into conductor ([[Agent-Skill]])
    # isolation: { mode: namespace }  # per-dispatch sandboxing ([[Isolation]])

triggers:
  github.pull_request:
    steps:
      - <<: *fixer                    # merge the anchor…
        id: fix                       # …then this step's own fields, which win
        prompt: "Resolve the conflict on {{.repo}}#{{.pr}}."

A step may of course carry these fields inline — an anchor is only how you avoid retyping them.

Two things to know:

  • Anchors are file-local. YAML resolves them per document, so an anchor defined in config.yaml is not visible in an imported conf.d/*.yaml. For reuse across files, use extends: on a map section (runtimes:/workflows:/handoffs:) or a named step, below. See Reuse.
  • An anchor copies fields, not identity. Two steps merging one anchor are still two identities. name: is the separate, deliberate opt-in to sharing one — see below.

Pointing at a step

There is no top-level steps: section. A step lives in the workflow or trigger that runs it, and anything that must point at a particular one addresses it there:

review-flow/architect     by the step's own id: (or name:, if it has no id:)
review-flow[2]            by position, for a step that carries neither

The container is a workflows: entry or a named trigger. That is the same slot the identity ladder uses below — id if present, else index — so a reference and an identity always name the same step by the same rule: review-flow/architect addresses exactly the step whose structural identity is workflow:review-flow/architect.

Two things use references:

  • a [[Teams|team:]] block, whose planner/worker/critic/reconcile roles each name a step. The role takes that step's behavior and its identity, so every team pointing at it shares one memory namespace, session pool, and track record;
  • a Packs overlay, where a consumer overrides one step of a pack by reference (packs.review.steps: { review-flow/review: { … } }).
workflows:
  roles:
    steps:
      - id: architect
        type: agent
        guidance: "Decompose before building. Name the risky part."
        prompt: "…"

triggers:
  - on: gh.issue_matched
    steps:
      - id: feature
        prompt: "Implement {{.url}}"
        team: { planner: roles/architect, worker: roles/architect }

The index form is the escape hatch, not the habit: it is positional, so inserting a step above shifts it. Give a step an id: and reference that.

agent: still parses, but it is now a free-form attribution label that selects nothing. What identifies a dispatch is its identity, below.

Step identity — the key everything defaults to

Memory scoping, session affinity, and outcome tracking all default their key to the step's identity. It must be stable across restarts, so it is a pure function of config — never a per-run value:

  1. an explicit name: — author-pinned, and shareable: two steps with the same name share one memory namespace, one session pool, and one track record — exactly the reuse a shared agent: fixer gave you. A team: role inherits the name of the step it references, for the same reason. Merging an anchor does not do this: an anchor copies fields and leaves identity alone.
  2. structural — the enclosing qualified trigger/workflow plus the step's slot: its id:, else its name:, else its index in the list — github.pull_request/security, or github.pull_request/2. This is the default, and it is the same slot a step reference uses.
  3. a deterministic fingerprint of the step's definition, for a step with no enclosing context.

Editing a step's prompt does not rotate its identity. Reordering may, for a step that has neither a name: nor an id: — give it one to pin it.

id: is not rung 1. It is the run-local handle steps.<id>.outputs.* addresses and it is near-universal, so treating it as a global identity would make two unrelated triggers that both use id: fix silently share a track record. Instead id: supplies the structural slot, which is what makes structural identity survive reordering.

Step forms

A step sets exactly one form — what it does:

Form Spelling
dispatch an agent type: agent (or a bare prompt: with agent:)
ask typed questions decide: + state: / questions: — answered with probabilities by a decision runtime or any agent runtime; see Decide-Steps
call a connector verb uses: <connector>.<verb> + options:
run code use: <engine> + code:, or use: cli + command: — see Code-Steps
run a program type: command + command:
call a workflow call: <workflow> + with: — see Workflows
a helper step sleep:, log:, set:, assert:, fail:, wait_for: — conductor's own flow control, below

Two of those spellings moved, and both old ones still parse:

  • use: selects the code ENGINE. It resolves like a connector's or a runtime's use: — the builtin (cli), then the official plugin repo's engines/<name>, which is where the scripting engines (js, go-embed, risor, lua) now live, then an explicit repo or path — and it also takes a host interpreter by name (bash, python3) or by path. run: is the same key and every run: js / run: bash config keeps working; set one or the other, not both.
  • call: is the workflow call. A step-level use: used to mean this. conductor config migrate (and the boot auto-migration) rewrites every use: <workflow> to call: <workflow>; workflow: remains valid as the older spelling of the same field.
steps:
  - { id: triage, use: js,  code: "return { sev: ctx.body.severity }" }
  - { id: build,  use: cli, command: [make, -C, ./svc, release] }
  - { id: review, call: review-flow, with: { pr: "{{.pr}}" } }

Helper steps

A helper step is a form conductor runs itself: no agent, no engine, no connector, no command behind it. There is nothing to dispatch, nothing to install and nothing to grant — which is the point, because the alternative was a use: cli step shelling out to a coreutil just to pause a flow.

The family, each a form in its own right:

Helper Does
sleep: <duration> Pause the flow for a positive duration.
log: <message> Render a message (templated) into the run log — a breadcrumb.
set: {<key>: <value>} Publish computed values as this step's outputs, read downstream as {{.<id>.<key>}}.
assert: <expr> Fail the run unless the condition is truthy — the same grammar as if:.
fail: <message> Stop the run with a rendered message; guard it with if: for a conditional abort.
wait_for: {…} Poll a read verb until a condition holds, or a timeout — below.

Each is mutually exclusive with type: / use: / uses: / call: and with the other helpers — a step does one thing — and each takes id:, if:, for_each: and step hooks like any other step, appears in the run timeline, and runs the same under [[One-Shot|conductor once]] as on the daemon.

sleep, log, set, assert, fail

steps:
  - log: "addressing {{.repo}}#{{.pr}}"          # a breadcrumb in the run log
  - { id: bail, if: "!run_id", fail: "no run_id on this event" }
  - assert: "checks_passed && !draft"            # stop unless the guard holds
  - id: vars
    set: { branch: "fix/{{.pr}}", attempts: 1 }  # {{.vars.branch}} downstream
  - sleep: 5s                                    # pause, cancellation-aware

What to know:

  • sleep's duration is the one every other field takes — 500ms, 5s, 30m, 7d, 1d12h (Configuration) — and must be positive; sleep: 0 and a negative are refused at load. Cancellation cuts it short: a shutdown, a step timeout:, or a budget cut-off leaves a sleep: 30m at once. A dry run prints [dry-run] would sleep 5s and moves on.
  • set values keep their type. attempts: 1 reads back as the number 1, not "1", and a value lifted from an earlier step's list/map arrives intact. Give the step an id: — that is how the values are addressed.
  • assert and fail both stop the run on failure, the same way any step error does (the run reports workflow_failed). assert is the conditional form — it fails unless the expr is truthy — and fail is unconditional, so it usually sits behind an if:. Both render/evaluate in a dry run too, the way if: does.
  • In an agent-authored plan each helper is a class an agent may only emit if the operator listed it in policy.agent_authored.verbs. They reach nothing outside conductor, but they still spend the run — admitting one stays a decision.

wait_for

Polls a read verb until its output satisfies a condition, or a timeout elapses. It is the principled form of a bare sleep: before a step that depends on state settling — instead of guessing how long a cancellation takes to land, wait for the run to actually reach a terminal state:

steps:
  - uses: gh.cancel_run
    options: { repo: "{{.repo}}", run_id: "{{.run_id}}" }
  - wait_for:
      uses: gh.get_run                            # the read verb to poll
      options: { repo: "{{.repo}}", run_id: "{{.run_id}}" }
      until: "status == 'completed'"              # condition over its outputs
      every: 10s                                  # poll interval (default 10s)
      timeout: 2m                                  # give up after this — required
    id: settle
  - uses: gh.rerun_run
    options: { repo: "{{.repo}}", run_id: "{{.run_id}}" }
  • until: reads the polled verb's outputs at the top level — a gh.get_run returns status, conclusion, … so until: "status == 'completed'" sees them directly, alongside every fact ({{.repo}}) and prior-step output. Same expression grammar as if:.
  • The verb's options: are re-templated each poll, exactly as a uses: step would render them.
  • On success the last poll's outputs become the step's outputs, so a later step reads {{.settle.conclusion}}. On timeout the step fails — soften a flow that should continue anyway with continue_on_error: true.
  • timeout: is required and positive; every: defaults to 10s. The whole wait is cancellation-aware — a shutdown drops out immediately, and the timeout is just another deadline on the same context. A dry run does one stubbed poll, says [dry-run] would poll …, and does not wait.

The family exists so the next helper is a small, predictable addition rather than a new subsystem. Anything that needs no agent, engine, verb or command to execute belongs here.

Fields

Field Meaning
name Pins the step's identity (see above). Shareable on purpose.
id The step's slot: how outputs are addressed (steps.<id>.outputs.*), what a step reference points at, and — absent a name: — the structural half of its identity.
use The code ENGINE this step runs on (run: is the same key). See Code-Steps.
command The argv for use: cli and for type: command — a list of words, or one string split on whitespace (quotes honored, no shell).
call A workflow to run as this step, with with: for its inputs. See Workflows.
sleep Pause the flow for a positive duration — a helper step, above.
log Render a message into the run log — a helper step.
set Publish computed values as the step's outputs ({{.<id>.<key>}}) — a helper step.
assert Fail the run unless the expr is truthy — a helper step.
fail Stop the run with a rendered message — a helper step.
wait_for Poll a read verb (uses:) until until: holds or timeout: — a helper step.
model Which model to run: a fleet name, a model id, a wildcard, an inline list, or { any, required }. Unset → the runtime's models.default:, then a bare launch. See Model-Selection.
runtime A runtimes.<name> entry to run on (default: the default: true runtime, else the built-in paseo). See Runtimes.
thinking / mode Runtime launch hints, passed through where the runtime supports them.
workspace local or worktree — the existing checkout, or a fresh git worktree.
expect_push Mark a fixer step that must LAND a change on the target. A clean run that leaves work unlanded (a non-empty proposed diff never pushed) becomes a no_progress failure instead of a silent success — see Workflows hooks. Leave off for a review/judge step.
wait_timeout How long a foreground dispatch waits before giving up.
archive_when_done Whether the agent (and its conductor-created workspace) is archived the moment its step finishes. Forced off for a background: hand-off step.
watch Reactive live step: run steps: every every: (a mini-workflow — fact reads + if-guarded actions). Actions: step.bail (reason gone → tear down), step.rerun (re-run this step on the new state), or a workflow: step (re-run a workflow — re-review). Works on ANY live step, not just hand-offs — e.g. a long-processing step can bail when its input disappears or rerun when new work arrives. (handoff.bail/handoff.rerun are accepted as deprecated aliases.) Operator-owned. See Hand-offs.
idle_timeout Release an interactive hand-off still open after this long (drop the hold, archive the agent + workspace). The backstop for one nobody closed; the agent calling step.done is the precise signal. Off unless set. See Hand-offs.
labels Extra key=value labels on the dispatched agent.
host A Hosts SSH target this step's runtime launches on, overriding the runtime's own.
guidance Tone/format that stacks on the scoped baseline ([[Policy|policy.guidance]]) rather than replacing it. A string, a list, or { replace: … }. See Reuse.
memory Opt into shared-memory injection (Memory). true uses the run's context keys; a map { scopes, tags, limit } names arbitrary keys. Also gates writing: only an opted-in step may harvest its output into shared memory.
session This step's own session pool: { key, idle_ttl, max_lifetime, end_on }, namespaced to its identity. A step with none joins the runtime's overall pool. See below.
skill Verb tools + the secret broker over the daemon socket: { verbs, secrets_via, allow_secrets, max_calls }. See Agent-Skill.
isolation Per-dispatch sandboxing. Wins over the runtime's own; needs a runtime conductor launches itself (not paseo). See Isolation.
outcome_feedback true appends this step's own track record to its guidance. See Outcomes.
outcome_key Override the track-record key (default: the identity), so several steps can pool one record.

Note there is no budget here: a budget caps execution cost on a backend, so it lives on the runtime (Cost-Accounting).

Behavior

  • Every step of a multi-step Workflows can differ — a common pattern is a cheap planner step triaging an issue and handing off to a stronger fixer only when the triage justifies it (if: "{{.evaluate.has_context}} == true").
  • workspace governs the checkout lifecycle (persistent checkout vs. a fresh worktree per dispatch); the step's own checkout: (checkout-pr | branch-off | none) governs what git state that checkout is put into. Independent knobs — a workspace: worktree step can still run checkout: none for triage.
  • workspace also takes an object, { isolation: local|worktree, pin: <name> }. pin: names a runtime workspace the step ALWAYS runs in — created on first use, reused by every later run — so a long-lived triage or chat agent finds its working directory as it left it. Without a pin, a checkout: none step gets a fresh workspace per dispatch that is archived when the run finishes. A pin cannot be combined with an explicit checkout: checkout-pr or branch-off (conductor validate rejects it): those give each dispatch its own worktree, which is the opposite of reusing one. With checkout: unset the strategy comes from the trigger, so the pin applies to the runs with no repo context and is ignored on the ones that get a worktree.
  • archive_when_done: true steps are still protected from premature cleanup: reclaim skips one paused on a permission prompt, and an agent can hold itself alive with a .paseo-hold marker in its worktree.
  • Which models a runtime can actually run is DISCOVERED, per runtime — see Model-Discovery. conductor validate reports a required: fleet nothing can satisfy.

Session affinity — one agent per (runtime, model, key)

By default each dispatch gets a fresh agent. A session: block binds a live session to the rendered key, so later events reach the same agent as a follow-up with full prior context.

A binding is (runtime, resolvedModel, key). Both extra dimensions are structural, not preferences: a live agent is one model on one runtime, and you can resume neither a paseo session on codex nor an opus step into a haiku session. A useful consequence — because a pack assigns a fleet per step, its model assignments partition affinity for free.

It is declarable at two scopes:

runtimes:
  paseo:
    session:                          # the OVERALL pool: one agent per key,
      key: "{{.repo}}#{{.pr}}"        #   shared by every step without its own
      idle_ttl: 12h
      end_on: [ gh._closed ]

x-templates:
  reviewer: &reviewer
    type: agent
    session:                          # this step's OWN pool, namespaced to its
      key: "{{.repo}}#{{.pr}}"        #   identity — so the same key string is
      end_on: [ gh._closed ]          #   still a distinct session

Resolution per dispatch: the step's session: wins, else the runtime's, else a fresh agent.

  • Serialized per key. At most one prompt in flight per session; concurrent same-key events queue — the group: one-run-per-key guarantee, extended across the session's life.
  • Durable. The binding persists in conductor's own state (affinity.json, beside audit/dedup); after a restart the session resumes via the runtime's native handle (paseo re-binds the agent id, ACP session/load). While bound, the agent is never reclaimed.
  • Eviction. Idle past idle_ttl, older than max_lifetime, or an end_on event (matched as <connector>.<kind>, a bare <kind>, or <type>.<kind>, rendered against the event's own context). github's PR close and merge both arrive as the internal _closed signal, so end_on: [ gh._closed ] is the evict-when-done form.
  • _closed is eviction-only. It matches end_on: but is not a trigger event — on: gh._closed fails validation.
  • Runtime support. Needs a session-persistent runtime — paseo or ACP. One-shot cli runtimes fall back to a fresh agent per event; pair with memory: (Memory) for continuity there.
  • vs group: and memory. group: batches a burst into one run; memory injects durable facts; session affinity reuses a live conversation. They compose.

Follow-up turns return Queued run refs: on paseo the prompt is delivered via paseo send (no captured output for later steps); an ACP follow-up returns the turn's output. A follow-up to a dead session is detected, unbound, and replaced by a fresh spawn.

Bindings written before the re-key carry no runtime/model and are ignored on restore, so the next event starts a fresh session. Sessions are short-lived by design (24h idle default), so nothing durable is lost — a track record would have been another matter, which is why that one is preserved explicitly.

Migration from agents:

conductor config migrate (and the automatic boot migration) decomposes each profile, and it is careful about one thing above all: your accumulated history carries over. Memory, sessions, and outcomes used to key off the agent NAME; they now key off the step IDENTITY. So the migration INLINES each profile's behavior onto every step that referenced it, carrying name: <the old agent name> — the identity ladder's top rung — onto each one.

Where a profile was referenced from several steps IN THE SAME FILE, those steps share a YAML anchor parked under x-migrated: instead of each getting a copy. Where the steps are in different files, each gets its own: anchors do not cross imports:, and duplicated config that works beats DRY config that does not parse. Because agents: commonly sat in config.yaml while the triggers naming it sat in conf.d/*.yaml, the profile table is gathered from the whole import tree before any file is rewritten.

A profile nothing referenced is dropped with a note — there is no top-level section left to park it in, and its behavior is in the .pre-connectors backup.

The keys therefore come out identical to what your box already has on disk: outcomeStats["fixer"] stays outcomeStats["fixer"], engagements still match, and memories written under the old agent:fixer scope are still recalled through a compatibility alias.

Old New
agents.<n> its fields, inlined on each referencing step, with name: <n>
provider + model model: (an exact pin — a migration never invents a fleet)
provider alone nothing — that named a backend, not a model, so it becomes a bare launch
budget runtimes.<the runtime it ran on>.budget
controller runtime
behavior fields the same key on each referencing step
agents.<child>.extends: <parent> flattened — the parent's fields are copied into the child before inlining
agent: <n> on a step the profile's fields, plus name: <n> (or <<: *n when shared in-file)

Anything the decomposition has no home for is dropped with a note in the migration summary, never silently.

Agent-driven workflows

An agent step's final output may carry a ```plan block — steps in the normal grammar that conductor validates, guards (policy.agent_authored, Policy), and runs deterministically. Failures route back to this agent's session (above) for a bounded revise loop, so a planning agent should keep sessions; pairing with memory: lets it recall what worked. The full loop — plan, choose from the catalog, supervise, promote — is in Workflows.

Explanation

A step and a runtime answer different questions. The step answers "what should run" — which model, what tone, what workspace lifecycle. The runtime answers "how is it run" — which process or API actually executes it. There is no provider:: a backend was never a property of the work, so a step with model: claude-opus-5 can run on the built-in paseo dispatcher, on agent-deck, or through opencode's HTTP API, unchanged, just by pointing runtime: at a different entry — the model still routes through whichever runtime is selected. The one place this decouples is an ACP or cli runtime: there, the runtime's own agent:/command: names the tool directly (e.g. gemini over ACP), so the step's model: has nothing to route and is ignored. With no runtimes: configured at all, every agent step runs on paseo, so this distinction is invisible until you actually introduce a second runtime. See Runtimes for the full resolution order and runtime kinds (Controllers is the legacy name).

The step lifecycle: step.done, step.bail, step.rerun

Every agent conductor launches carries the built-in step connector's verbs (the handoff.* spellings remain as deprecated aliases):

  • step.done — the agent's own completion signal, auto-granted to every dispatch. Calling it releases the agent and its conductor-created workspace (ledger-gated: an agent can only ever release itself — the target comes from its token, never an argument). On a step with an output_schema, the call ALSO delivers the result: conductor call step.done --output '<result>' is one atomic final action, validated against the schema at the verb boundary (an invalid value gets a precise, retryable error). Any JSON value shape the schema admits works — object, array, string, number, boolean, enum.
  • step.bail / step.rerun — watch: rule actions on the LIVE step (not agent-callable): tear it down when its reason is gone, or supersede it with a re-run on the current state.

How agents learn about it: schema steps get the delivery instruction as their schema directive; other foreground steps get a short "run step.done when finished" footer; hand-offs get theirs in the hand-off guidance. Where the signal is REACHABLE: paseo agents call the conductor CLI (creds injected in env; remote hosts via the SSH tunnel); ACP and opencode runtimes get the same surface as native MCP tools (step_done). A bare cli one-shot runtime has no skill surface — there the step boundary itself is the completion signal (conductor is blocked on the process), schema output rides the reply, and the footer tells the agent to just finish normally if the command is unavailable. Reclaim never depends on the call alone: a foreground step is archived at its step boundary regardless, and hand-offs have idle_timeout as the backstop.

Clone this wiki locally