DSH | dsh-turn-doctor | Find out which timeout actually killed a failed turn #5135
d3vmeh
started this conversation in
Show Your Plugins!
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Project URL:
https://github.com/d3vmeh/dsh-turn-doctor
npm:
https://www.npmjs.com/package/dsh-turn-doctor
Introduction
This came out of a problem that's been mentioned a few times in DSH discussions.
Back in #3157, someone put it pretty well:
That's basically what
dsh-turn-doctordoes.A request can fail in several different places: DSH's idle watchdog (
streamIdleTimeoutMs), the SDK request timer (timeoutMs), Node/undici's HTTP timers, the model server itself, or an individual tool call.The annoying part is that the error shown in the UI often isn't enough to tell those apart.
Request timed out.can come from more than one timer.terminatedmight mean undici stopped waiting for the response body, or it might mean the model server died.So you can end up increasing one timeout, retrying, and still having the request die at exactly 5:00. That's also come up in #4518.
dsh-turn-doctorputs its own stopwatch around each model request and tracks things like time to first byte, longest period of silence, and total request time. When something fails, it combines those measurements with the actual error and tries to identify which layer caused it.For example:
It also catches a couple of failures I wasn't getting much visibility into before, especially failed compactions and tool timeouts. In my own logs I found six compaction summaries that had quietly hit their token cap.
You can type
/whyin chat to see the recent diagnoses for that session, including failures from subagents.How it integrates with DSH
It's a host-side Cordis plugin.
It observes the
llm/streamwaterfall to collect timing data without modifying the chunk stream, plusagent/request-errorand a few session events such asllm/retry,turn/end,compaction/end, andtool/result.It doesn't retry requests or change any DSH behavior. It's just watching and reporting.
Verdicts currently show up in:
/whyturn-doctor/verdictevents in the session logThat last one should also make it possible to add a proper row for these in the web UI later.
I don't duplicate retry countdowns since DSH already renders those.
Install:
It works without any config and assumes DSH's default timeout values.
If you've changed your route timeouts, you can tell the plugin what they are so it can make better timing-based diagnoses:
A few notes
dsh-turn-doctoris more focused on separating the different timeout layers using timing data, and it also covers compactions and tool calls.This is also the fourth plugin in the little local-model setup I've been building around DSH:
dsh-llm-gate(DSH | dsh-llm-gate | One request at a time for local llama.cpp, no more subagent timeouts #4995) keeps concurrent requests within the model server's available slotsdsh-context-budget(DSH | dsh-context-budget | Keep a local model's context at a size your GPU handles well #5078) keeps context usage under controldsh-fetch-timeouts(DSH | dsh-fetch-timeouts | Stop slow local models (Ollama, LM Studio) from being cut off at 5 minutes #5124) lets you get past Node's 5-minute HTTP cutoffsdsh-turn-doctortries to tell you what actually went wrong when a turn still failsMIT licensed. Would appreciate feedback, especially if anyone has failure cases it misclassifies or doesn't recognize yet.
> Unofficial project, independently developed and maintained by community members.
All reactions