Trace Compare & Live Maze — see how your agent actually explored (main path, detours, backtracks, subagent branches) #3638
Replies: 2 comments 1 reply
|
Drawing the maze from the session log (including subagent branches) without a second LLM judge is the useful part — transcripts hide the backtracks. Compare-two-logs on the same axis is a real debugging loop, not a screenshot gallery. (#3582 is the duplicate; this thread is the one to keep.) If you want it in the community catalog: https://github.com/ylwl1997/dshbase/issues/new?template=plugin-submission.yml |
|
Thanks — you nailed the two design bets: verdicts are deterministic and drawn from the session log alone (no second model, so every ✗ is reproducible and explainable on hover), and same-axis compare is exactly how I debug (two providers, same task). And thanks for the quick verification over at dshbase — impressively fast pipeline. Appreciate the dedupe note on #3582 too. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Every agent session is a maze run: the model commits to a path, hits dead ends, backtracks, retries. The transcript shows you what it said — not where it wandered. I built a plugin that draws the wandering.
Two surfaces:
Honesty rules I care about: verdicts are deterministic and explainable (no LLM judging) — error flags → failure signatures (head/tail windows only, so error text quoted in a git log doesn't count as a failure) → per-tool-class rules, plus behavioral blind-retry detection borrowed from AgentLens-style waste analysis. Every verdict shows its rationale on hover. Tokens are read from
assistant/messageusage, not stream-chunk counts. Idle gaps fold; everything else stays wall-clock.New in v0.5.0: the whole UI is bilingual (en/中文) and live-follows your dsh language setting — verdicts are stored as language-neutral structured keys, so switching languages re-renders instantly without re-parsing.
Already listed in awesome-deepseek-harness (Visualization) and Awesome DSH Plugins.
Install (rc.6–rc.8 verified, tracked per rc):
Repo (docs in EN/中文): https://github.com/lamost423/dsh-trace-compare
Feedback very welcome — especially session logs where the verdict rules get something wrong; the thresholds live in one
VERDICT_RULESconstant and I calibrate against real corpora.一句话中文版:把智能体真实的探索过程画出来——主干、失败支路、折返点、子代理支路,全部落在同一根墙钟时间轴上;支持双会话同轴对比与会话内实时生长。仓库 README 即中文文档(English
All reactions