Skip to content

Restarting a seat

joshdaugherty edited this page Sep 25, 2026 · 2 revisions

Restarting a seat

Describes robot-council/cli v0.4.24 and robot-council/core main as of 2026-09-25.

A restart is a new fleet session

Every robot-council mcp process starts unjoined, unless it was started with --auto-join, and joining starts a new fleet session (#127). Nothing carries a session from one bridge process to the next. So restarting a harness (and so its bridge) ends one session and, once the agent joins again, starts another.

What ending a session does

When the bridge stops on SIGTERM or SIGINT, it:

  1. ends its fleet session, which core records as session.gone with the reason ended;
  2. clears its fleet events from the sink, keeping only its own bridge. records, and gives up after two seconds if another process holds the sink's lock (v0.4.3, #228).

What the session held is released on core's next presence sweep, not at the moment it ends. Core has no listener releasing work on session.gone; instead, every sweep runs release steps that return a gone session's claimed, in-progress or blocked tasks to pending and release its locks (robot-council/core RobotCouncilServiceProvider::registerReleases()). Core schedules the sweep every minute by default (schedule.sweep_sessions), so a fleet on the defaults releases the work within about a minute.

Where PHP has no pcntl (PHP on Windows ships without it), the bridge cannot catch a signal, so a signal kills it before it ends its session, which the sweep then marks stale and gone on its thresholds (by default 5 and 30 minutes). A bridge whose harness closes its stdin still ends its session, signal or not. This is read from source and has not been measured.

Resuming a conversation

These traps were met restarting live seats on 2026-09-24 so their channel could be enabled (#242, #242):

  • A session in the VS Code extension cannot be switched over in place. claude --resume in a terminal starts a separate process, and the extension keeps its old connection, without the channel. Close the conversation in VS Code first, then resume it in a terminal.
  • Resume from the conversation's own project directory, usually the repository's primary checkout rather than a worktree. From anywhere else, --resume only prints to resume, run: cd ….
  • Resume by conversation id, with the flag: cd <project directory> && claude --resume <conversation-id> --dangerously-load-development-channels server:robot-council. The id is the sessionId in the bridge's MCP log.
  • Check registration per conversation: a Channel notifications registered line carrying that conversation's own id, after the restart (Starting a seat).
  • Leaving the VS Code copy open leaves two copies running. The feed showed no new session.joined for the resumed seats, and later work arrived under their original session ids, so neither the feed nor task_list shows which copy acted.

Does resuming keep the fleet session? Unmeasured.

One report (#242) found that after a resume the agent was offered no join tool, presence_heartbeat returned the original session as active, and a task placed before the restart was started afterwards. That does not fit how the bridge works: a new bridge process offers only join until the agent calls it, unless it was started with --auto-join. The later finding that the VS Code copies were still running (#242) explains it better: the resumed conversation's calls were most likely reaching the old bridge, whose session was still open. If so, the resumed copy's own bridge was never joined, and could not receive notices.

Until this is measured with the old copy confirmed closed, treat a resumed conversation as needing to join again, and confirm with presence_heartbeat which session it is on.

Two bridges in one checkout share a sink

The sink is keyed by service, harness and project, and, when no project is named, by the checkout the session was launched in (v0.4.23, #299). Before v0.4.23, every session launched from one user-level configuration shared one sink, whatever checkout it was in. Now two running bridges share a file only when they were launched in the same checkout, or given the same --project. An exiting bridge clears its fleet events, and with them its live sibling's (#228); the sibling also stops repeating its notice early (#231). The "two copies running" trap above is exactly this case, since both copies run in the conversation's own project directory.

Resuming from a different folder changes the sink. A bridge started in another checkout reads and writes that checkout's sink, so events already waiting for the old one stay there. Resume from the conversation's own project directory, as above.

Clone this wiki locally