Skip to content

Releases: Zen1th53/marshal

MARSHAL v0.0.5-rc.5

MARSHAL v0.0.5-rc.5 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 01 Oct 21:03
e85f1d5

MARSHAL v0.0.5-rc.5 — Commands that act

A release candidate. It is cut from the integration branch
release/v0.0.5-rc.1, not from main: rc.4's tree plus #75, which merges the
command sweep (#47) and the implementation stages #48–#74. None of them is
merged into main yet.

Since rc.4

Live testing of rc.4 found commands that were listed in /help and completion
but only printed that they were unavailable. They now act, or they are gone:

  • Operator commands are authenticated locally. The operator is the OS user
    that owns the project's state directory. Agents never receive that identity;
    the socket still carries only the worker protocol, checked by peer UID.
  • Goals and tasks. /goal edits the goal with a version check and reports
    its revision history. Approvals are decided with /approve and /reject.
    /tasks shows project-scoped tasks with their lifecycle and scope, and
    /msg posts to the existing team session.
  • Run control. /pause, /resume run:<id> and /cancel act on runs of
    this session. Budget limits from the goal contract stop dispatch when spent.
    Model and effort profiles apply to new work.
  • Checks. The alignment guard (scope lock, blast radius, goal drift) runs on
    finished tasks in advisory mode and its decisions are recorded. /blind is
    hidden until blind interpretations are collected. /fingerprint recomputes
    failure signatures; evidence, diagnostics, /runtime and /store read back
    real state.
  • Restore. Checkpoints can be listed, inspected and restored with
    verification. /backup restore restores the database from a live session
    only when no process holds it, so no write-ahead log is lost (Linux).
  • Providers. Provider commands follow an explicit grammar, and a typo runs
    nothing. Operations are qualified against the installed CLI version.
    /sessions separates native conversations from governed runs, and --last
    picks this project's newest conversation instead of the provider's global
    one.
  • /apply no longer guesses. It needs a Codex task ID, snapshots the
    project first and reports the files that actually changed.
  • /diff takes staged, unstaged and untracked scopes; skills show where
    they come from.

In rc.4

Found in live testing of rc.3, all in #46:

  • Enter respects the completion you picked. Choosing settings from the
    /marshal menu with the arrows or Tab and pressing Enter used to open a
    Marshal session. Enter now accepts the highlighted candidate; with nothing
    picked, it still runs a finished command such as /goal .
  • /marshal on its own no longer opens a session. It shows the Marshal's
    status and the command list. /marshal chat opens the conversation.
  • A mistyped subcommand runs nothing. /marshal setings answers "Did you
    mean /marshal settings?" instead of starting a planning run with the typo as
    its goal. A goal of more than one word still starts planning.
  • Malformed commands are refused. /marshal status now or
    /memory inject auto extra show the usage instead of acting.
  • The /marshal menu lists every subcommand, chat first.
  • /memory inject knows all four agents. It reports, previews and
    confirms the channel for Claude, Codex, OpenCode and Antigravity (agy), and
    works, like /memory peers, without a database.
  • No hint for what you cannot use. The activity panel and /help show the
    navigation shortcuts only when navigation is open to the session, and the
    Tab and Enter help matches what the keys do.

In rc.3

Found in live testing of rc.2, all in #40:

  • The Marshal no longer prints its protocol. Every agent reads its briefing
    where it reads instructions:

    • Claude from the system prompt;
    • Codex from developer instructions;
    • OpenCode and Agy from a per-session directory under .marshal/briefing/.

    You see the MARSHAL wordmark, a short "Begin." and the Marshal's
    introduction.

  • Claude as the Marshal works. Its protocol was dropped whenever the
    cross-agent memory reached the system prompt first. A second briefing now
    joins the first.

  • MARSHAL stays out of your repository.

    • Briefings are no longer written into the project's AGENTS.md or
      CLAUDE.md.
    • A block an earlier version left there is removed on the next launch.
    • The briefing directory ignores itself in git and is deleted when the
      session ends.
  • The Marshal introduces itself. It states its model, that it is MARSHAL's
    Marshal here, and whether the run is Standard or ULTRA. It opens with some
    swagger, then stays plain.

  • The protocol follows the agreed business process end to end.

    • Each planned task says why its worker was chosen and what it should
      produce.
    • The plan table shows the control level.
    • The final report lists every criterion as verified or not tested, the
      budget spent, remaining risks and what was not done.
  • A MARSHAL wordmark plays while the Marshal starts, a different one each
    time. The launch line and the Marshal panel no longer name the provider
    behind the Marshal.

Install it deliberately. install.sh and /update follow the latest release,
and a candidate is not that.

What changed

One model plans with you, then marshals the work to other agents.

/marshal chat opens a conversation with the appointed Marshal model (#24). It
follows a fixed protocol, compiled into the binary and pinned by a digest (#31,
#35):

  1. it introduces itself;
  2. it asks which language to work in;
  3. it asks your goal;
  4. it reads the current state of the project before asking anything else;
  5. it asks what the plan depends on, one question at a time, each with a
    recommendation.

It records nothing you did not answer.

What you agree stays written down (#35). When the Marshal is done, it writes
a plan pack: REQUIREMENTS.md, 00_INDEX.md and one tasks/<id>.md per task.
After /exit, MARSHAL moves the pack to .marshal/marshal/runs/<run>/plan/ and
shows where it is. You can read it and correct it. /marshal approve binds the
pack as it stands into the approval, and every worker's brief carries its own
task note and the requirements.

Workers run the approved plan, and nothing else.

  • After approval, each task runs in its own worktree, and MARSHAL re-runs the
    task's checks itself (#24).
  • Dependent tasks start from the merged work they depend on. A reassigned
    worker starts fresh from the task's base (#28).
  • Choose the control level before approval (#29):
    • /marshal settings control strict: every task carries instructions that
      the worker follows exactly;
    • free: the worker chooses its approach within the task's files.
  • In user and hybrid acceptance modes, /marshal return <task> <reason> sends
    work back (#27).
  • Every reviewing role runs in a session of its own, so a single installed
    provider is enough, ULTRA included (#25).

ULTRA Execution is switched on in the TUI (#38). The
MARSHAL_ULTRA_EXECUTION variable is gone:

  • every session starts with execution off;
  • /ultra start turns execution on when this installation holds a verified
    entitlement;
  • /ultra stop asks first, and /ultra stop confirm turns execution off.

Commands outside the TUI, such as marshal goal, always ask you.

/ultra says more (#36, #37):

  • it shows when the grant ends, separately from the five-minute session lease;
  • a refusal from the Community Cloud shows the server's reason instead of a
    bare status code.

One Cloud identity per machine (#30). The installation identity now lives in
~/.config/marshal/cloud_state.json instead of in every project, and MARSHAL
reports its real version to the Cloud.

The Ctrl+N navigation surface is closed (#39). It still has screens with no
capability behind them, and no end-to-end run has shown that it works. Until
that is proven, Ctrl+N and Esc refuse for every session. Everything remains
available from the composer.

Known limits

  • Alignment is advisory. The guard records its decisions but does not yet
    block a task.
  • Provider qualification is narrow. Only Codex 0.159.2, Claude 2.1.286,
    OpenCode 1.18.16 and agy 1.2.7 are qualified; other versions keep their
    commands as unqualified pass-through.
  • Live database restore is Linux-only. Elsewhere, restore offline.
  • Not run with real models end to end. Codex, Claude and Agy have each
    been started as the Marshal and followed the protocol's opening. None has
    yet taken a plan from /marshal chat through approval to accepted work.
  • OpenCode is a worker, not a Marshal model. The Marshal's review and
    verification turns need schema-enforced output, which OpenCode does not
    offer. OpenCode receives its briefings, but a small local model may not
    follow them.
  • Keep the window open for a run. MARSHAL has no background service, so the
    window must stay open while a run is in progress.
  • Older identity carried over. On first start, a machine that already had a
    project-level Cloud identity adopts it. Check /ultra if your entitlement
    seems to be missing.

Verification

  • The release gate (scripts/community-release-gate.sh) passed on this commit
    before tagging.
  • The release workflow publishes the artifacts with checksums, an SBOM, a
    release manifest and build provenance.

MARSHAL v0.0.5-rc.4

MARSHAL v0.0.5-rc.4 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 30 Sep 14:31

MARSHAL v0.0.5-rc.4 — Marshal mode

A release candidate. It is cut from an integration branch, not from main: the
branch merges the open pull requests #24–#31 and #35–#39 onto v0.0.4, in their
stack order, then #40 and #46, and none of them is merged into main yet.

Since rc.3

Found in live testing of rc.3, all in #46:

  • Enter respects the completion you picked. Choosing settings from the
    /marshal menu with the arrows or Tab and pressing Enter used to open a
    Marshal session. Enter now accepts the highlighted candidate; with nothing
    picked, it still runs a finished command such as /goal .
  • /marshal on its own no longer opens a session. It shows the Marshal's
    status and the command list. /marshal chat opens the conversation.
  • A mistyped subcommand runs nothing. /marshal setings answers "Did you
    mean /marshal settings?" instead of starting a planning run with the typo as
    its goal. A goal of more than one word still starts planning.
  • Malformed commands are refused. /marshal status now or
    /memory inject auto extra show the usage instead of acting.
  • The /marshal menu lists every subcommand, chat first.
  • /memory inject knows all four agents. It reports, previews and
    confirms the channel for Claude, Codex, OpenCode and Antigravity (agy), and
    works, like /memory peers, without a database.
  • No hint for what you cannot use. The activity panel and /help show the
    navigation shortcuts only when navigation is open to the session, and the
    Tab and Enter help matches what the keys do.

In rc.3

Found in live testing of rc.2, all in #40:

  • The Marshal no longer prints its protocol. Every agent reads its briefing
    where it reads instructions:

    • Claude from the system prompt;
    • Codex from developer instructions;
    • OpenCode and Agy from a per-session directory under .marshal/briefing/.

    You see the MARSHAL wordmark, a short "Begin." and the Marshal's
    introduction.

  • Claude as the Marshal works. Its protocol was dropped whenever the
    cross-agent memory reached the system prompt first. A second briefing now
    joins the first.

  • MARSHAL stays out of your repository.

    • Briefings are no longer written into the project's AGENTS.md or
      CLAUDE.md.
    • A block an earlier version left there is removed on the next launch.
    • The briefing directory ignores itself in git and is deleted when the
      session ends.
  • The Marshal introduces itself. It states its model, that it is MARSHAL's
    Marshal here, and whether the run is Standard or ULTRA. It opens with some
    swagger, then stays plain.

  • The protocol follows the agreed business process end to end.

    • Each planned task says why its worker was chosen and what it should
      produce.
    • The plan table shows the control level.
    • The final report lists every criterion as verified or not tested, the
      budget spent, remaining risks and what was not done.
  • A MARSHAL wordmark plays while the Marshal starts, a different one each
    time. The launch line and the Marshal panel no longer name the provider
    behind the Marshal.

Install it deliberately. install.sh and /update follow the latest release,
and a candidate is not that.

What changed

One model plans with you, then marshals the work to other agents.

/marshal chat opens a conversation with the appointed Marshal model (#24). It
follows a fixed protocol, compiled into the binary and pinned by a digest (#31,
#35):

  1. it introduces itself;
  2. it asks which language to work in;
  3. it asks your goal;
  4. it reads the current state of the project before asking anything else;
  5. it asks what the plan depends on, one question at a time, each with a
    recommendation.

It records nothing you did not answer.

What you agree stays written down (#35). When the Marshal is done, it writes
a plan pack: REQUIREMENTS.md, 00_INDEX.md and one tasks/<id>.md per task.
After /exit, MARSHAL moves the pack to .marshal/marshal/runs/<run>/plan/ and
shows where it is. You can read it and correct it. /marshal approve binds the
pack as it stands into the approval, and every worker's brief carries its own
task note and the requirements.

Workers run the approved plan, and nothing else.

  • After approval, each task runs in its own worktree, and MARSHAL re-runs the
    task's checks itself (#24).
  • Dependent tasks start from the merged work they depend on. A reassigned
    worker starts fresh from the task's base (#28).
  • Choose the control level before approval (#29):
    • /marshal settings control strict: every task carries instructions that
      the worker follows exactly;
    • free: the worker chooses its approach within the task's files.
  • In user and hybrid acceptance modes, /marshal return <task> <reason> sends
    work back (#27).
  • Every reviewing role runs in a session of its own, so a single installed
    provider is enough, ULTRA included (#25).

ULTRA Execution is switched on in the TUI (#38). The
MARSHAL_ULTRA_EXECUTION variable is gone:

  • every session starts with execution off;
  • /ultra start turns execution on when this installation holds a verified
    entitlement;
  • /ultra stop asks first, and /ultra stop confirm turns execution off.

Commands outside the TUI, such as marshal goal, always ask you.

/ultra says more (#36, #37):

  • it shows when the grant ends, separately from the five-minute session lease;
  • a refusal from the Community Cloud shows the server's reason instead of a
    bare status code.

One Cloud identity per machine (#30). The installation identity now lives in
~/.config/marshal/cloud_state.json instead of in every project, and MARSHAL
reports its real version to the Cloud.

The Ctrl+N navigation surface is closed (#39). It still has screens with no
capability behind them, and no end-to-end run has shown that it works. Until
that is proven, Ctrl+N and Esc refuse for every session. Everything remains
available from the composer.

Known limits

  • Not run with real models end to end. Codex, Claude and Agy have each
    been started as the Marshal and followed the protocol's opening. None has
    yet taken a plan from /marshal chat through approval to accepted work.
  • OpenCode is a worker, not a Marshal model. The Marshal's review and
    verification turns need schema-enforced output, which OpenCode does not
    offer. OpenCode receives its briefings, but a small local model may not
    follow them.
  • Keep the window open for a run. MARSHAL has no background service, so the
    window must stay open while a run is in progress.
  • Older identity carried over. On first start, a machine that already had a
    project-level Cloud identity adopts it. Check /ultra if your entitlement
    seems to be missing.

Verification

  • The release gate (scripts/community-release-gate.sh) passed on this commit
    before tagging.
  • The release workflow publishes the artifacts with checksums, an SBOM, a
    release manifest and build provenance.

MARSHAL v0.0.5-rc.3

MARSHAL v0.0.5-rc.3 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 30 Sep 06:27

MARSHAL v0.0.5-rc.3 — Marshal mode

A release candidate. It is cut from an integration branch, not from main: the
branch merges the open pull requests #24–#31 and #35–#39 onto v0.0.4, in their
stack order, then #40, and none of them is merged into main yet.

Since rc.2

Found in live testing of rc.2, all in #40:

  • The Marshal no longer prints its protocol. Every agent reads its briefing
    where it reads instructions:

    • Claude from the system prompt;
    • Codex from developer instructions;
    • OpenCode and Agy from a per-session directory under .marshal/briefing/.

    You see the MARSHAL wordmark, a short "Begin." and the Marshal's
    introduction.

  • Claude as the Marshal works. Its protocol was dropped whenever the
    cross-agent memory reached the system prompt first. A second briefing now
    joins the first.

  • MARSHAL stays out of your repository.

    • Briefings are no longer written into the project's AGENTS.md or
      CLAUDE.md.
    • A block an earlier version left there is removed on the next launch.
    • The briefing directory ignores itself in git and is deleted when the
      session ends.
  • The Marshal introduces itself. It states its model, that it is MARSHAL's
    Marshal here, and whether the run is Standard or ULTRA. It opens with some
    swagger, then stays plain.

  • The protocol follows the agreed business process end to end.

    • Each planned task says why its worker was chosen and what it should
      produce.
    • The plan table shows the control level.
    • The final report lists every criterion as verified or not tested, the
      budget spent, remaining risks and what was not done.
  • A MARSHAL wordmark plays while the Marshal starts, a different one each
    time. The launch line and the Marshal panel no longer name the provider
    behind the Marshal.

Install it deliberately. install.sh and /update follow the latest release,
and a candidate is not that.

What changed

One model plans with you, then marshals the work to other agents.

/marshal chat opens a conversation with the appointed Marshal model (#24). It
follows a fixed protocol, compiled into the binary and pinned by a digest (#31,
#35):

  1. it introduces itself;
  2. it asks which language to work in;
  3. it asks your goal;
  4. it reads the current state of the project before asking anything else;
  5. it asks what the plan depends on, one question at a time, each with a
    recommendation.

It records nothing you did not answer.

What you agree stays written down (#35). When the Marshal is done, it writes
a plan pack: REQUIREMENTS.md, 00_INDEX.md and one tasks/<id>.md per task.
After /exit, MARSHAL moves the pack to .marshal/marshal/runs/<run>/plan/ and
shows where it is. You can read it and correct it. /marshal approve binds the
pack as it stands into the approval, and every worker's brief carries its own
task note and the requirements.

Workers run the approved plan, and nothing else.

  • After approval, each task runs in its own worktree, and MARSHAL re-runs the
    task's checks itself (#24).
  • Dependent tasks start from the merged work they depend on. A reassigned
    worker starts fresh from the task's base (#28).
  • Choose the control level before approval (#29):
    • /marshal settings control strict: every task carries instructions that
      the worker follows exactly;
    • free: the worker chooses its approach within the task's files.
  • In user and hybrid acceptance modes, /marshal return <task> <reason> sends
    work back (#27).
  • Every reviewing role runs in a session of its own, so a single installed
    provider is enough, ULTRA included (#25).

ULTRA Execution is switched on in the TUI (#38). The
MARSHAL_ULTRA_EXECUTION variable is gone:

  • every session starts with execution off;
  • /ultra start turns execution on when this installation holds a verified
    entitlement;
  • /ultra stop asks first, and /ultra stop confirm turns execution off.

Commands outside the TUI, such as marshal goal, always ask you.

/ultra says more (#36, #37):

  • it shows when the grant ends, separately from the five-minute session lease;
  • a refusal from the Community Cloud shows the server's reason instead of a
    bare status code.

One Cloud identity per machine (#30). The installation identity now lives in
~/.config/marshal/cloud_state.json instead of in every project, and MARSHAL
reports its real version to the Cloud.

The Ctrl+N navigation surface is closed (#39). It still has screens with no
capability behind them, and no end-to-end run has shown that it works. Until
that is proven, Ctrl+N and Esc refuse for every session. Everything remains
available from the composer.

Known limits

  • Not run with real models end to end. Codex, Claude and Agy have each
    been started as the Marshal and followed the protocol's opening. None has
    yet taken a plan from /marshal chat through approval to accepted work.
  • OpenCode is a worker, not a Marshal model. The Marshal's review and
    verification turns need schema-enforced output, which OpenCode does not
    offer. OpenCode receives its briefings, but a small local model may not
    follow them.
  • Keep the window open for a run. MARSHAL has no background service, so the
    window must stay open while a run is in progress.
  • Older identity carried over. On first start, a machine that already had a
    project-level Cloud identity adopts it. Check /ultra if your entitlement
    seems to be missing.

Verification

  • The release gate (scripts/community-release-gate.sh) passed on this commit
    before tagging.
  • The release workflow publishes the artifacts with checksums, an SBOM, a
    release manifest and build provenance.

MARSHAL v0.0.5-rc.2

MARSHAL v0.0.5-rc.2 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 29 Sep 19:43

MARSHAL v0.0.5-rc.2 — Marshal mode

A release candidate. It is cut from an integration branch, not from main: the
branch merges the open pull requests #24–#31 and #35–#39 onto v0.0.4, in their
stack order, and none of them is merged into main yet.

v0.0.5-rc.1 was tagged but never published. Its release gate failed on
TestMarshalWiredUltraHasIndependentVerifier: the ULTRA verifier check still
refused the Marshal's own provider, so ULTRA verification failed on a machine
with a single provider. #25 now fixes that, and this candidate includes the fix.

Install it deliberately. install.sh and /update follow the latest release,
and a candidate is not that.

What changed

One model plans with you, then marshals the work to other agents.

/marshal chat opens a conversation with the appointed Marshal model (#24). It
follows a fixed protocol, compiled into the binary and pinned by a digest (#31,
#35):

  1. it introduces itself;
  2. it asks which language to work in;
  3. it asks your goal;
  4. it reads the current state of the project before asking anything else;
  5. it asks what the plan depends on, one question at a time, each with a
    recommendation.

It records nothing you did not answer.

What you agree stays written down (#35). When the Marshal is done, it writes
a plan pack: REQUIREMENTS.md, 00_INDEX.md and one tasks/<id>.md per task.
After /exit, MARSHAL moves the pack to .marshal/marshal/runs/<run>/plan/ and
shows where it is. You can read it and correct it. /marshal approve binds the
pack as it stands into the approval, and every worker's brief carries its own
task note and the requirements.

Workers run the approved plan, and nothing else.

  • After approval, each task runs in its own worktree, and MARSHAL re-runs the
    task's checks itself (#24).
  • Dependent tasks start from the merged work they depend on. A reassigned
    worker starts fresh from the task's base (#28).
  • Choose the control level before approval (#29):
    • /marshal settings control strict: every task carries instructions that
      the worker follows exactly;
    • free: the worker chooses its approach within the task's files.
  • In user and hybrid acceptance modes, /marshal return <task> <reason> sends
    work back (#27).
  • Every reviewing role runs in a session of its own, so a single installed
    provider is enough, ULTRA included (#25).

ULTRA Execution is switched on in the TUI (#38). The
MARSHAL_ULTRA_EXECUTION variable is gone:

  • every session starts with execution off;
  • /ultra start turns execution on when this installation holds a verified
    entitlement;
  • /ultra stop asks first, and /ultra stop confirm turns execution off.

Commands outside the TUI, such as marshal goal, always ask you.

/ultra says more (#36, #37):

  • it shows when the grant ends, separately from the five-minute session lease;
  • a refusal from the Community Cloud shows the server's reason instead of a
    bare status code.

One Cloud identity per machine (#30). The installation identity now lives in
~/.config/marshal/cloud_state.json instead of in every project, and MARSHAL
reports its real version to the Cloud.

The Ctrl+N navigation surface is closed (#39). It still has screens with no
capability behind them, and no end-to-end run has shown that it works. Until
that is proven, Ctrl+N and Esc refuse for every session. Everything remains
available from the composer.

Known limits

  • Not run with real models end to end. No live Claude, Codex or Agy session
    has yet taken a plan from /marshal chat through approval to accepted work.
    The unit, integration and PTY tests pass.
  • Keep the window open for a run. MARSHAL has no background service, so the
    window must stay open while a run is in progress.
  • Older identity carried over. On first start, a machine that already had a
    project-level Cloud identity adopts it. Check /ultra if your entitlement
    seems to be missing.

Verification

  • The release gate (scripts/community-release-gate.sh) passed on this commit
    before tagging.
  • The release workflow publishes the artifacts with checksums, an SBOM, a
    release manifest and build provenance.

MARSHAL v0.0.4

Choose a tag to compare

@github-actions github-actions released this 22 Sep 17:15
dc41db0

MARSHAL v0.0.4 — One shared channel

Agents in a project now work from one ordered channel instead of from a
briefing that went stale the moment it was written. Each agent drops what it
does into it as it works, and each reads its own view of it — filtered by what
you decided before the work started.

The channel

Every agent contributes as it works. Claude Code and Codex write
append-only history that MARSHAL reads as it grows. OpenCode and Antigravity
keep SQLite, opened read-only and never written to, which SQLite serves to
readers while a writer holds the file. No agent waits until it exits.

An agent joining late joins the conversation. It used to start from empty,
on the reasoning that its briefing had already summarised what came before. A
summary is not the work. A reader now resumes from its own cursor, so an agent
opening while another is five steps into a task sees those five steps.

An agent that was closed catches up. Entries accumulate whether or not the
recipient is running. Each carries the time it happened, and a boundary marks
where a session begins, so nothing older reads as though it just arrived. A
full view drops its oldest entries rather than sealing itself, because the
recent work is the part a returning agent needs.

.marshal/stream/events.jsonl   the channel — ordered, append-only
.marshal/stream/cursors.json   how far each agent has looked
.marshal/inbox/<agent>.md      that agent's view of it
.marshal/live-peers            who joins, and who each agent sees

Choosing who sees whom

Two decisions, both made before the work starts, with /memory peers or in
.marshal/live-peers:

participants: claude, codex, opencode, agy
agy: all                  # every other agent
claude: all
codex: opencode, agy      # not claude
opencode: none            # contributes, reads nothing

Joining and seeing are separate. An agent can contribute while reading almost
nothing, and that is an arrangement rather than a gap: models differ in what
they can use.
A capable one does better seeing everything the others did; a
smaller one does worse, because context it cannot follow is context it can be
confused by. The two directions between any pair may disagree — a reviewer can
read the implementer without the implementer reading the reviewer.

An agent is never shown its own work, and that is not a setting. It already
knows what it did.

Capture reports itself

Memory capture was always live; nothing said so, because the only report came
after the agent exited. The native CLI owns the terminal for a whole session,
so MARSHAL cannot draw a counter — it writes one instead, to
.marshal/<agent>/live-status.json: records imported, entries shown, last sync
and any capture error.

Workspace

  • Completion works on arguments. The menu used to close after a command and
    a space, exactly where you had most reason to expect it, and Tab then took the
    first candidate silently. It now opens for a command's arguments too, and Tab
    steps through the candidates the way a shell does, leaving the menu up.
  • Enter settles, a second Enter runs. Settling on /memory is where you
    reach for Tab again to complete peers; running on the same keystroke ran
    something half-written.
  • Antigravity is Agy cli in the menu, after the command you actually type.
  • Gemini leaves the provider menu. Its adapter, its doctor probe and
    Import Gemini JSONL are unchanged — this narrows what the TUI offers, not
    what MARSHAL can run.

Known limits

Delivery is a pull. A running CLI owns its own input; MARSHAL cannot
interrupt it. The view is a file that is current whenever the agent looks, and
the briefing says so rather than implying the agent is kept in sync.

Verification

Tested against the installed CLIs on 2026-09-22 — Claude Code 2.1.278, Codex
0.155.1, OpenCode 1.18.16, agy 1.2.7 — with real sessions rather than
doubles:

  • Each agent was run and asked to emit a marker. The watcher read all four real
    stores and put 22 entries into the channel, with all four markers present.
  • Each agent was then asked to read its own view. All four correctly named the
    other three and quoted their markers. None reported its own.
  • With codex: none set, a real Codex session was run inside the real TUI.
    The channel went from 22 entries to 64 with that session's marker in it, while
    codex's own view stayed at its header — none of the other three agents'
    markers reached it. The data was there and it was withheld.

The arrangement space is tested exhaustively rather than by example: 65,536
configurations written, read back and checked to still mean the same thing for
all sixteen author/reader pairs, and every filter a reader can have rendered to
a real file and read back.

Also in this release

Internal, with no change in behaviour: the CI gate no longer runs the whole
suite twice and is split by cost; the store suite migrates once into a template
instead of 311 times, taking it from 400 seconds to 184 under the race
detector; and the PTY waits scale with the race detector, which had been
failing on a loaded runner and reading as a broken TUI rather than a slow one.

MARSHAL v0.0.4-rc.1

MARSHAL v0.0.4-rc.1 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 22 Sep 11:41

MARSHAL v0.0.4-rc.1 — One shared channel

A release candidate. It carries the cross-agent channel and nothing else: it is
cut from the branch that work lives on, not from main, so the four other
changes open at the time are not in it.

Install it deliberately. install.sh and /update follow the latest release,
and a candidate is not that.

What changed

Every agent drops its work into one ordered channel, as it happens.

Cross-agent exchange used to be point to point: a running session copied what it
did into every other agent's mailbox. Four agents meant four copies of one
event, a guard against writing it twice, and a table of who posts to whom —
bookkeeping the shape created rather than the problem.

There is one channel now. What each agent sees is decided when it reads: a
filter naming the authors it takes, and a cursor saying how far it has looked.

.marshal/stream/events.jsonl   the channel — ordered, append-only
.marshal/stream/cursors.json   how far each agent has looked
.marshal/inbox/<agent>.md      that agent's view of it
.marshal/live-peers            who joins, and who each agent sees

An agent joining late joins the conversation. It used to start from empty,
on the reasoning that its launch briefing had already summarised what came
before. A summary is not the work. A reader now drains from its cursor, so an
agent opening while another is five steps into a task sees those five steps.

An agent that was closed still catches up. Entries accumulate whether or not
the recipient is running. Each carries the time it happened, and a boundary
marks where a session begins, so nothing older reads as though it just arrived.

All four agents now reach the channel while they work. Claude Code and Codex
write append-only history. Antigravity and OpenCode keep SQLite, opened
read-only and never written to; SQLite serves readers under WAL while a writer
holds the file. OpenCode was the last exception and it was an inconsistency
rather than a limit — its private store is no more private than agy's, which was
already being read.

Reading a store MARSHAL does not own is a coupling, and it is guarded. OpenCode's
shape is checked before each read; if it has moved, that path stands down and
the supported CLI export takes over, which cannot run mid-session and so
delivers at exit — later than it should be, never wrong. What is read is
assembled into the same document the export produces and handed to the same
decoder, so the live path cannot select different fields. The model's hidden
reasoning is excluded there, once, for both
— checked against 199 real
reasoning blocks, none of which reached a transcript.

Choosing who sees whom

Two decisions, in .marshal/live-peers or through /memory peers:

participants: claude, codex, opencode, agy
agy: all                  # every other agent
claude: all
codex: opencode, agy      # not claude
opencode: none            # contributes, reads nothing

Joining and seeing are separate. An agent can contribute while reading almost
nothing, and that is an arrangement rather than a gap: models differ in what
they can use. A capable one does better seeing everything the others did; a
smaller one does worse, because context it cannot follow is context it can be
confused by.

An agent is never shown its own work, and that is not a setting. It already
knows what it did.

Capture reports itself

The native CLI owns the terminal for a whole session, so MARSHAL cannot draw a
counter — it writes one, to .marshal/<agent>/live-status.json: records
imported, entries shown, last sync, and any capture error. Memory capture was
always live; nothing said so.

Fixes in the workspace

  • A command that changes the arrangement now reprints it and marks the row that
    moved, instead of answering with one line and leaving the change off screen.
  • The channel report no longer lists an author that never joined, which was a
    promise the channel could not keep.
  • Colour marks state; every agent name is the same colour. Column padding is
    computed on the visible text rather than the escape bytes.
  • Autocomplete: /memory offered two of its five subcommands, argument
    completion offered the parent command's list, and an empty word offered
    nothing. The longest registered prefix now wins, and a trailing space lists
    what the command accepts.

Known limits

Delivery is a pull. A running CLI owns its own input; MARSHAL cannot
interrupt it. The view is a file that is current whenever the agent looks, and
the briefing says so rather than implying the agent is kept in sync.

This candidate is not main. It does not contain the Gemini provider-surface
change, the CI split, the store test template, or the PTY timing fix.

Verification

  • go test ./internal/tui/ — the full suite, including the PTY conformance
    tests, passing.
  • 65,536 channel arrangements written, read back, and checked to still mean the
    same thing for all sixteen author/reader pairs; every filter a reader can have
    rendered to a real file and read back.
  • Verified against the installed CLIs on 2026-09-22 — Claude Code 2.1.278, Codex
    0.155.1, OpenCode 1.18.16, agy 1.2.7 — by decoding their real transcripts
    with the production watcher.

MARSHAL v0.0.3

Choose a tag to compare

@github-actions github-actions released this 18 Sep 11:24
44a672f

MARSHAL v0.0.3 — Native Antigravity, Guided Setup and Verified Updates

v0.0.3 adds Antigravity's CLI as a native session alongside Codex, Claude and
OpenCode, takes an empty directory to a ready project from one command, and lets
MARSHAL tell you when a newer release exists and install it on request.

Highlights

  • Native Antigravity sessions. Open the Antigravity CLI (agy) with
    marshal agy, /agy or F12, using your own configuration and sign-in.
    /agy continue, /agy resume <conversation>, /agy cli <args> and
    /agy <prompt> map onto agy's own flags. When agy exits, its visible
    conversation, tool calls and command output are saved to project memory;
    the model's reasoning is never read. Its work reaches the other agents'
    briefings, and theirs reaches agy through AGENTS.md. The Team panel now
    finds agy instead of reporting it unavailable.
  • Guided setup. marshal setup now offers the blocking steps it reports,
    in the order they have to happen: initialize a Git repository, make an empty
    first commit as a baseline, and set up MARSHAL for the project. Each step is
    asked for by itself and runs only on a yes. marshal init no longer has to be
    typed separately.
  • Update notice in the workspace. When a newer release is published, the
    activity panel says so, with the key that installs it:
    Update MARSHAL v0.0.4 is available [F10] Download and install · /update.
  • /update and marshal update. Check for a newer release from the
    workspace or the shell, and install it with /update install,
    marshal update install or F10.
  • Security and reliability improvements.
    • Updates are installed only after the archive matches the release's
      published SHA-256. A download that fails verification is not installed,
      and the binary in place is left untouched.
    • The new binary is put in place by a rename within its directory, so it is
      never half-written.
    • Checking and installing are separate. The workspace checks the release
      feed on its own, but installs nothing until you press F10 or run the
      install command, and F10 installs only the release the notice is showing.
    • The release source is fixed to this repository and cannot be redirected.
    • Downloads are bounded by silence, not by size. marshal update install
      abandons a transfer only after 30 seconds without a byte, so a slow link
      that keeps delivering completes; install.sh gives every request a
      connect timeout, abandons a transfer below 1 KB/s for 30 seconds, and
      retries it, instead of waiting forever on a stalled connection.
    • setup changes nothing without an answer: where its output is not going
      to a terminal, and for setup status, it only reports. MARSHAL's automatic
      repair set is unchanged.
    • agy's conversation databases are opened read-only. Its storage format is
      not published, so only fields observed to carry visible conversation and
      tool evidence are read, and a step that cannot be parsed is skipped rather
      than guessed at. Only conversations this launch created or continued are
      imported, and one recorded against another workspace is left to it.
    • The first commit setup makes is empty, so no file in the directory is
      added to the repository without your decision.
    • Git's "Author identity unknown" and its revision-parsing error on a
      repository with no commits are now reported in plain terms with the step
      to take.
    • Handoff checkpoints and rollbacks are ordered by time rather than by the
      text of their timestamps. Trimmed fractional seconds made "…:07Z" sort
      after "…:07.5Z", which could return the wrong checkpoint as the latest and
      made a store test fail intermittently.

Configuration

  • MARSHAL_NO_UPDATE_CHECK=1 stops MARSHAL contacting the release feed at all.

Installation

Install the latest checksum-verified Linux release:

curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh | sh

Pin this release explicitly:

MARSHAL_VERSION=v0.0.3 \
  sh -c "$(curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh)"

From v0.0.3 on, marshal update install upgrades an existing installation.

Published assets include Linux amd64 and arm64 archives, SHA-256 checksums, an
SPDX SBOM, a release manifest and GitHub build-provenance attestations.

Verification

  • Antigravity: decoding of user input, visible answers, tool calls, command
    output and failures with reasoning excluded; workspace attribution; a
    real-terminal session with memory capture and F12; and an end-to-end run
    with agy 1.2.5 whose conversation was found in MARSHAL memory after exit
  • Setup: real-terminal tests for each offered step, a declined step, a run with
    no terminal, and setup status
  • Update: verified install, refused install on a checksum mismatch, version
    comparison, the opt-out variable, a slow but steady download that outlasts
    the feed timeout, and a stalled download abandoned with the binary untouched
  • Update against the published feed, and a real install from v0.0.1 to v0.0.2
  • Workspace notice, F10 behaviour and command registration tests
  • Checkpoint ordering across timestamps that differ only in trimmed fractions
  • Full internal package and TUI test suites
  • Release workflow build, test, race, vulnerability, conformance, clean-install
    and manifest gates before publication

MARSHAL v0.0.2

Choose a tag to compare

@github-actions github-actions released this 17 Sep 08:07
2a448c6

MARSHAL v0.0.2 — Native OpenCode and Automatic Session Memory

v0.0.2 adds OpenCode as a first-class native terminal session alongside Codex
and Claude. It also corrects terminal cursor placement for the ❯ composer
prompt and hardens native-session memory capture.

Highlights

  • Launch OpenCode directly with marshal opencode, F9, or /opencode.
  • Start, continue, resume and fork OpenCode sessions from the MARSHAL TUI.
  • Pass native OpenCode arguments through without shell re-parsing.
  • Automatically import visible OpenCode conversation and bounded tool evidence
    into project candidate memory when the native process exits.
  • Exclude private reasoning, snapshots and provider metadata from memory.
  • Preserve normal memory secret scanning before any imported record is stored.
  • Feed bounded cross-agent project context and live inbox information into
    native OpenCode sessions.

Fixed

  • The composer keeps the visible ❯ separator while Ctrl+Left places the
    hardware cursor on the first input character.
  • OpenCode export capture preserves visible conversation instead of persisting
    redacted placeholders.
  • Transient partial OpenCode exports retry once.
  • Existing OpenCode history is baselined before launch, so exit capture imports
    the session created or updated by that run.
  • Stable session record IDs make repeated imports idempotent when provider
    metadata changes.
  • Oversized provider JSONL events no longer prevent later conversation from
    being imported.
  • Repeated peer-history rejection diagnostics appear once instead of filling
    the native-session result pane.

Installation

Install the latest checksum-verified Linux release:

curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh | sh

Pin this release explicitly:

MARSHAL_VERSION=v0.0.2 \
  sh -c "$(curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh)"

Published assets include Linux amd64 and arm64 archives, SHA-256 checksums, an
SPDX SBOM, a release manifest and GitHub build-provenance attestations.

Self-hosted development

This release was developed in MARSHAL's own native-agent workspace. MARSHAL
coordinated Codex and OpenCode sessions in the MARSHAL repository and captured
completed native-session context into its canonical project memory. The final
archives are still produced independently by the pinned release workflow from
the exact annotated tag.

Verification

  • Native OpenCode unit and real-PTY lifecycle tests
  • Codex and Claude native-session regression tests
  • Session importer, idempotency and secret-boundary tests
  • Full internal package and TUI test suites during development
  • Release workflow build, test, race, vulnerability, conformance, clean-install
    and manifest gates before publication

MARSHAL v0.0.1

Choose a tag to compare

@github-actions github-actions released this 16 Sep 09:16

MARSHAL v1.0.1 — Canonical Community Consolidation and Production Hardening

v1.0.1 is a reconciliation and hardening release. It consolidates valid work
from overlapping branches onto the current Community runtime, removes false
production claims, and makes the existing release path deterministic and
verifiable.

Highlights

  • Canonical Community runtime and SQLite memory schema v72
  • Governed memory consolidation, task-change cursors, retrieval quality gates,
    and current session importers
  • Read-only Community Resource Awareness with bounded local/Ollama probes
  • Fail-closed production Web boundary for fixture-only handlers
  • Self-contained, symlink-safe clean initialization
  • Pinned release dependencies/actions and synchronized licensing notices
  • Rewritten user documentation grounded in v1.0.1 behavior
  • Reproducible Linux amd64/arm64 archives, SPDX SBOM, checksums, release
    manifest, and GitHub build-provenance attestation

Fixed

  • doctor no longer executes optional provider probes unless
    --probe-providers is supplied.
  • Historical note: the former marshal web serve behavior and its fixture
    routes belonged to the pre-Community/Enterprise-split product. Community
    no longer ships a Web command, routes, or Web runtime.
  • Legal source evidence reads runtime_implementation_version and
    pack_version from committed blobs instead of stale defaults.
  • Clean marshal init can create the required policy/version defaults without
    a source checkout and rejects symlink replacements.
  • Repeated provider capability checks no longer collide in the durable audit
    stream; each decision retains fail-closed audit enforcement.
  • Concurrent role-binding revocations retry bounded SQLite lock contention and
    converge to one durable winner plus one conflict.
  • Network-required provider runs now return NET_ENFORCEMENT_UNAVAILABLE
    instead of opening a Bubblewrap network namespace that could bypass the
    endpoint allowlist proxy.
  • The release toolchain now uses Go 1.25.13, which clears the reachable
    standard-library vulnerabilities reported against Go 1.25.0.

Changed

  • Community resource reporting includes CPU/cgroup, RAM/swap, storage,
    accelerator, thermal, and loopback Ollama awareness. Recommendations remain
    advisory and do not change runtime concurrency or policy.
  • Community/Enterprise boundaries now explicitly exclude adaptive resource
    governors, fleet placement, and autonomous provider/model tuning.
  • README and operator docs distinguish implemented adapters, local probes, and
    authenticated E2E verification.

Verification

The mandatory Community release gate ran on 2026-08-26 with Go 1.25.13.

Gate Result
Go build / vet / test / race PASS
govulncheck ./... PASS — no vulnerabilities found
Sandbox, policy, authz, memory, migration, resource, provider, backup tests PASS
Web npm ci / typecheck / lint / 116 tests / build / embedded-asset parity PASS
Python pack conformance, release tooling, and legacy tooling PASS
Clean install, initialization, doctor, daemon, first local workflow, backup, Web start, restart persistence PASS
Source pack manifest PASS
Reproducible archives, SPDX SBOM, checksums, and release-manifest tests PASS

ESLint emitted one non-fatal no-useless-escape warning in
web/src/api/errors.ts; lint still exited successfully. Credentialed provider
qualification is reported separately below and is not included in the
mandatory gate PASS.

Provider verification

Provider path Adapter/probe Adapter/model E2E Canonical Runtime E2E
Codex PASS — local codex-cli 0.149.1 NOT_RUN — credentialed execution not enabled NOT_RUN
OpenCode + DeepSeek V4 PASS — workstation OpenCode 1.18.16 PASS — Flash 7.01s and Pro 6.64s; strict proof rerun also PASS NOT_RUN — endpoint-enforcing provider egress unavailable
OpenCode + Ollama PASS — workstation Ollama 0.32.9 and model inventory; release host service unavailable FAIL — qwythos-9b and blackarch-ai wrote incorrect proof content; qwen2.5-coder-abliterate:14b created no proof file NOT_RUN — endpoint-enforcing provider egress unavailable
Gemini CLI NOT_AVAILABLE — optional binary absent NOT_RUN — binary and credentialed execution unavailable NOT_RUN
Claude Code PASS — local Claude Code 2.1.198 NOT_RUN — credentialed execution not enabled NOT_RUN

The OpenCode model qualification ran on commit b37b187 in an isolated
temporary repository. Process exit zero was insufficient: the test required
the requested proof file and exact content. No local-model failure is reported
as PASS.

Known limitations

  • Endpoint-specific provider egress fails closed because Bubblewrap alone
    cannot enforce host/port allowlists and no enforcing proxy is wired.
  • Bubblewrap is the supported production sandbox backend; process-only fallback
    is explicit and limited to eligible R0/R1 work.
  • Fixture-only Web panels are disabled for live runtimes.
  • No independent third-party security audit is claimed.