Skip to content

[Bug] A provider-backed agent stays "online" with a live Shutdown button after it has shut down — remote liveness is derived from backend_agent_id, not presence (violates I3) #4730

Description

@wierdbytes

Summary

Pressing Shutdown on a provider-backed (remote) managed agent stops the agent, but the desktop keeps rendering it as online / running indefinitely, and the primary action button stays "Shutdown" forever — it never flips back to "Deploy".

The agent is not at fault: it publishes both of the signals the spec asks for (kind:10100 roster entry with "status":"offline", and kind:20001 presence "offline"), and the relay accepts and applies both. The desktop simply does not consult either one when it decides whether a remote agent is alive.

Root cause: for backend != Local, build_managed_agent_summary() derives the status exclusively from record.backend_agent_id.is_some(), and backend_agent_id is never cleared outside of agent deletion. Every liveness-shaped UI decision downstream is keyed off that one field.

This contradicts invariant I3 ("Presence is the status") in docs/remote-agents.md:200, which states that the deployment axis "is bookkeeping, not liveness", and it is not among the 8 entries in §Known Defects.

Environment

  • Buzz Desktop 0.5.4 (desktop/src-tauri/tauri.conf.json), macOS
  • Agent deployed through a third-party backend provider implementing the documented provider protocol (buzz-backend-*, local-host substrate)
  • All file:line references below verified against feccf4eabc23fdba94ce3537a194357ed17b197c

Steps to reproduce

  1. Create a managed agent with a provider backend and press Deploy. The agent comes up, appears online, and answers in a channel.
  2. Open the agent's profile panel and press Shutdown. A toast appears: "Shutdown command sent. Agent will stop shortly."
  3. The agent receives the owner !shutdown, drains, publishes its farewell (kind:10100 status: "offline"), publishes presence offline (kind:20001), and exits. It stops answering in the channel.
  4. Wait — minutes, hours, a restart of the desktop.

Expected

The agent's card / profile stops showing a live indicator, and the primary action becomes Deploy so the agent can be brought back.

Actual

  • The agent still renders with a green "online" dot.
  • The Runtime tab still shows a green Running dot.
  • The status badge still reads a green deployed.
  • The primary action button still reads Shutdown and is still styled as live.
  • Pressing Shutdown again succeeds (green toast, !shutdown re-sent into a channel of a process that no longer exists) — because the offline roster entry still carries channel_ids, so resolveManagedAgentChannelId still resolves.
  • There is no way to redeploy the agent from this surface. handleAgentPrimaryAction (desktop/src/features/profile/ui/useAgentLifecycleActions.ts:30) branches on isManagedAgentActive, which is permanently true, so the startManagedAgentWithRules arm is unreachable.

Evidence: the agent reported correctly, the desktop ignored it

Agent-side log at shutdown:

13:38:41.122Z  info  draining              reason=owner_shutdown
13:38:43.518Z  warn  relay connection closed

Process gone. I then reconnected to the relay with the agent key (NIP-42 + NIP-OA auth tag) and queried its own events:

=== kind 10100 (roster) ===
created_at: 1785850721  → 2026-08-04T13:38:41Z   (exactly the drain instant)
content: {"name":"…","agent_type":"agent","status":"offline","capabilities":[],
          "channels":[],"channel_ids":["650ac09e-…","81f2eb5e-…"]}

=== kind 20001 (presence) ===
(none — the relay holds no presence for this pubkey)

Both axes are correct on the relay:

  • the roster entry was replaced with "status":"offline";
  • presence was cleared, which is exactly what crates/buzz-relay/src/handlers/event.rs:823-830 does on a kind:20001 whose content is "offline".

So the desktop had everything it needed. It just never looks.

Root cause

1. Remote status is a function of one persisted field.

desktop/src-tauri/src/managed_agents/runtime.rs:148-169:

let (status, pid, log_path) = if record.backend != BackendKind::Local {
    // Two-axis status model for remote agents:
    //   …
    //   Live axis (relay presence, polled by frontend): online/away/offline.
    //   Shown as a PresenceDot next to the agent name. …
    // After !shutdown the agent goes offline (presence) but stays "deployed"
    // (infrastructure still exists). This is intentional …
    let status = if record.backend_agent_id.is_some() {
        "deployed".to_string()
    } else {
        "not_deployed".to_string()
    };
    (status, None, String::new())
} else { /* local: real pid probe */ };

The comment's two-axis model is a reasonable design. The bug is that the frontend never implements the second axis for the decisions that matter — it treats the control-plane axis as liveness.

2. backend_agent_id is write-once. Set at desktop/src-tauri/src/commands/agents.rs:501 on a successful deploy; the only paths that clear it are record creation defaults and deletion. There is no undeploy (deferred to "v2" per the comment at runtime.rs:163), and stop_managed_agent explicitly refuses remote agents (desktop/src-tauri/src/commands/agents.rs:1244-1248). So status is pinned at "deployed" for the entire life of the record.

3. Every liveness-shaped UI decision keys off it.

Surface Location Behaviour
Active predicate desktop/src/features/agents/lib/managedAgentControlActions.ts:34 status === "running" || "deployed" → permanently true
Primary action label .../managedAgentControlActions.ts:38-47 provider + active → "Shutdown", never "Deploy"
Button "live" styling desktop/src/features/profile/ui/UserProfilePanelSections.tsx:364-367 agentActionLive = same predicate, inlined
Green online dot desktop/src/features/agents/ui/AgentRuntimeAvatarControl.tsx:130 <PresenceDot status="online" />hardcoded literal, gated only on isActive. Presence is never read here
Status badge desktop/src/features/agents/ui/AgentStatusBadge.tsx:27-31 green deployed; the "Starting…" de-escalation branch tests status === "running", so it can never fire for a remote agent
Runtime tab dot desktop/src/features/profile/ui/UserProfilePanelSections.tsx:129-148 deployed"running"bg-emerald-500

ManagedAgentRow.tsx:230 and ProfileHero (UserProfilePanelSections.tsx:485-497) do render a real presence-driven PresenceDot — so the app disagrees with itself on the same screen once the two axes diverge.

Contributing (secondary) issue: nothing is refetched after !shutdown

handleStop (desktop/src/features/agents/ui/useManagedAgentActions.ts:240-267) invalidates and refetches nothing for the provider branch, unlike handleStart (:220-221). The same is true of useMembersSidebarActions.ts:173-188 and useAgentLifecycleActions.ts:30-41.

Meanwhile useRelayAgentsQuery polls kind:10100 on a 5-minute interval (desktop/src/features/agents/hooks.ts:337), and usePresenceQuery backstop-polls at 60s (desktop/src/features/presence/hooks.ts:94).

So even the surfaces that do honour presence lag by up to a minute, and roster-derived status by up to five — right at the moment the user is watching for feedback on an action they just took. Fixing the root cause without also invalidating ["presence"] / relay-agents / managed-agents after a successful !shutdown send would leave a visible multi-minute lie.

Why this is a spec violation, not a design choice

docs/remote-agents.md:200-206, invariant I3:

(I3) Presence is the status. D derives a remote agent's live state exclusively from relay presence events self-signed by the agent key: online/away/offline (kind:20001, ephemeral, WS-published). The deployment axis (deployed/not_deployed, from the stored backend_agent_id) is bookkeeping, not liveness.

The implementation inverts this for every decision listed in the table above. This is not in §Known Defects (docs/remote-agents.md:1572-1681, defects 1–8).

I3 also promises a bounded wrong dot (180s PRESENCE_TTL_SECS). Today the wrong dot is unbounded: it survives shutdown, app restart, and machine reboot, because it is not a presence dot at all.

Suggested fix

Smallest correct change is to give remote agents the live axis the comment already describes, and let it drive the "is this thing alive" decisions:

  1. Introduce a live signal for remote agents. Either
    (a) have build_managed_agent_summary() fold relay presence into the returned status for backend != Local, or
    (b) keep the Rust summary purely control-plane and add a frontend helper — e.g. isManagedAgentLive(agent, presenceLookup) — that returns presence !== "offline" && presence !== undefined for provider agents and falls back to status for local ones.
    (b) is less invasive and keeps deployed/not_deployed honest as bookkeeping, per I3.
  2. Route the liveness-shaped call sites through it: getManagedAgentPrimaryActionLabel, agentActionLive, AgentRuntimeAvatarControl (drop the hardcoded status="online" and pass the real presence through), AgentStatusBadge, resolveRuntimeTabStatus. Keep isManagedAgentActive for the genuinely control-plane questions (e.g. "should Delete warn about orphaning a remote deployment?" in deleteManagedAgentWithRules, which already reads presenceLookup and gets this right).
  3. Invalidate after a !shutdown send["presence"], relay-agents, managed-agents — in stopManagedAgentWithRules' callers, so the transition is visible immediately rather than on the next 5-minute poll.
  4. Once (1)–(3) land, an offline remote agent's button becomes Deploy, and startManagedAgentWithRulesdeploy reaches the provider, which is already a converge-to-one-live-instance operation (§Deploy State Machine) and safely re-adopts or recreates. That closes the "cannot bring it back" half of this report without needing the v2 undeploy.

Related

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions