Follow-ups queued during an active turn are never admitted after the turn ends (non-cancel wake-latch variant) #4952
Ashmurthy64
started this conversation in
General
Replies: 1 comment
|
I turned this into a source-backed Agent runtime runbook: https://github.com/sandbaseai/deepseek-harness-handbook/blob/main/docs/en/troubleshooting/queued-followup-wake-latch.md — it defines the expected event ordering, isolates the post-turn wake latch, and covers idempotence, restart recovery, cancellation, and safe containment. Published in v0.5.369: https://github.com/sandbaseai/deepseek-harness-handbook/releases/tag/v0.5.369 |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Title: Follow-ups queued during an active turn are never admitted after the turn ends (non-cancel wake-latch variant)
Version: dsh v0.1.0-rc.8 (commit 141eb6f), web profile, per-user instances
Summary: User messages submitted while a turn is running are spliced into the agent inbox as queued follow-ups, but after the active turn completes (
turn/end), the driver never wakes to admit them. The agent status remainsidle, the messages stay parked forever, and the UI shows the prompt accepted (input clears) with no response rendered. Every subsequent retry piles up the same way.Repro (web profile):
Expected: the queued follow-up opens its own turn after the first completes (
followup()semantics — "the item becomes the sole ordinary message of its own turn").Actual: no
turn/startever fires for the queued message. Session event log showsuser/message+agent/inbox/splicedwith no subsequent turn events. The session appears dead; only a new session recovers.Evidence (persisted session logs,
session.jsonl.zstd,zstdcat):In the fully-traced small session, the parked messages were spliced at event index 6 — during the first active turn — and the log ends at
turn/end(index 41) with the queue never drained. The uniform "2 parked" across sessions is the natural signature: one stall plus one retry.Analysis: This looks like the non-cancel sibling of the cancel-convergence wake-latch bug fixed 2026-08-07 (
.agents/notes/implemented/bug-fix/2026-08-07-cancel-convergence-wake-latch.md). There, a wake submitted during cancel convergence could miss its turn boundary; here, a follow-up queued during an ordinary active turn loses its post-turn wake. Relevant semantics inpackages/core/agent/src/runtime-types.ts(send/followup/whenIdle) —followup()promises the queued item becomes its own turn, but the wake that should claim it afterturn/endappears not to fire.Impact: any user who sends while a turn is running loses the message silently. For interactive fleet deployments this is a top-tier defect — sessions feel dead after the first occurrence, because every subsequent input parks the same way.
Environment: Ubuntu 24.04 arm64, Node 22, instance-per-user web profile, sessions persisted via session-persistence-jsonl. Model route claude-fable-5 via the Anthropic catalog provider (repro is model-independent — the failure is pre-model; no LLM request is ever issued for the parked messages).
All reactions