You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Hi all — I've been using DSH Web as my main working environment for about a month, and I want to give back the three problems that cost me the most time. Each one comes with measurements, and I've separated what I measured from what I merely suspect.
TL;DR — (1) Long-session cost didn't disappear, it turned into manual work I now maintain by hand. (2) One patch-level upgrade of DSH took six steps, including elevation and downtime, because plugin compatibility has no contract. (3) In structural-analysis tasks the model is overconfident, and its negative conclusions ("there is no X") are the expensive ones, because nobody re-checks a "no".
DSH Web, profile=web, on Windows; long-lived runtime; multiple third-party plugins; a mix of Web GUI sessions and IM (WeCom) group sessions feeding the same session.
Daily driver for real delivery work, not experiments: my last project was a self-contained 1.6 MB operations dashboard built, verified and handed over entirely through DSH sessions.
I also maintain a public MIT-licensed DSH skill package.
Scope warning: this is one machine, one very heavy usage pattern (Windows + many plugins + sandbox + scheduled tasks). The numbers are order-of-magnitude for my profile, not averages, and the environment is more extreme than most users'. Items marked [hypothesis] are not verified.
1. Long sessions: the cost didn't disappear, it turned into manual work
What I measured
Cost structure of one heavy session:
Metric
Value
Model calls
584
Tokens
~274M
Cost
≈ ¥31.66
Spent re-reading context
98.6%
Model output
0.17%
So cost ≈ average context size × number of calls, and is almost independent of how much is actually produced.
The part I think matters more than the ratio
To work around the context ceiling, I built an external memory system by hand:
Artifact
Size
Handover doc A (runtime / plugins / upgrades)
605 lines / 66 KB
Handover doc B (business project handover)
1,229 lines / 111 KB
Gotchas I had to write down as "never do this again"
31
Session-history search tool
self-written (no official one available)
Plugin × version compatibility health check
self-written
The context cost didn't go away — it became the hours I spend maintaining 187 KB of handover documents. And it still isn't enough.
Long-session side effects, same session
Metric
Value
Session log
4.69 MB compressed / 16.68 MB uncompressed
Records
3,549 (tool/call 710, tool/result 720)
UI must render
644 chat bubbles + 1,430 tool cards
Largest records
43 > 50 KB, max 107 KB
Web process
599 MB RSS, 710 s CPU
Result: the UI visibly degrades, and the user is effectively pushed into starting a new session (losing continuity) or living with the lag.
The session log is not searchable
The session log is appended multi-frame zstd. A one-shot decompress returns only the first frame (the file looks empty); streaming decompression dies on frame two with Unknown frame descriptor. What I ended up doing is scanning for the magic number 28 B5 2F FD, splitting frames and decompressing them one by one: 4.69 MB → 1,946 frames → 16.68 MB → 3,549 JSONL records. Until you reverse that, session history is a black box for both the user and the model.
Suggestions
Make the context budget visible and archivable in layers. What needs to survive a long task is conclusions, contracts, gotchas — not the raw transcript. Compressing "finished work" into a searchable, verifiable form would remove most of the re-reading. (I'm already doing this by hand; that's the signal.)
Ship the model-facing session-query tool in the default composition. Verified today on 0.1.7-rc.2 / profile=web: dsh-session-query and dsh-session-query-sqlite (FTS5 backend) are installed, but the model-facing package dsh-tool-session-query is not. Consequence: in a long session the model has no built-in way to look up "what did the user actually ask earlier" — it either relies on what's still in context, or writes its own script to decompress the session log (which is what I do).
Document and expose the session format ("one zstd frame per append batch") and provide an official read path, e.g. dsh session grep <pattern>.
Deal with long-session rendering: built-in turn collapsing / message virtualisation, plus a threshold hint ("this session has 3,000+ records"). Note that a community plugin already exists that collapses completed turns into a single line — so this is a shared pain, not my edge case.
Treat handover summaries as a product feature. Heavy users are hand-maintaining 187 KB of them. Auto-generate, verify completeness, update incrementally.
Clarify IM ↔ Web session binding. Measured: one Web session and several WeCom group sessions share the same session id, so group messages and GUI conversation land in the same history (16.68 MB within ~3 hours) and their contexts contaminate each other. Also, IM bindings do not follow when you fork a session. Is this intended, or a default worth revisiting?
2. Plugin compatibility: one patch-level upgrade took six steps
What I measured
Three plugin compatibility failures in a single patch-level upgrade, each printing dsh: skipping profile bundle on every start — and note this only goes to stderr, never to .out, so grep in the normal log finds nothing:
A third-party plugin declared its peer whitelist only up to 0.1.6-alpha.2 → the plugin had to be upgraded before DSH would load it.
A local link: plugin whose source had been moved away, leaving a broken symlink in node_modules → had to be removed from the manifest.
An official composition bundle that no longer exists in the new install → had to be removed from bundles.
An upstream icon-vocabulary rename broke a plugin, and the error is unreadable. The new harness renamed ui-primitives icons as a whole (Icon*Outline16 → Icon*OutlineRegular), so named imports in the plugin became undefined and the panel crashed on render. All the user sees is:
That error tells you nothing about which plugin or which API changed; I had to go read the plugin author's release notes to find out. Worse, the same root cause reports a different number across versions (#138 → #130), which reads like a new problem.
Version semantics are counter-intuitive, so upgrades silently do nothing. The manifest said ^0.53.3. Under 0.x semantics that pins the minor line (>=0.53.3 <0.54.0), so 0.57.0 can never be installed — install reports success and the version does not move. I only found out because the panel kept crashing.
Upgrading requires downtime and elevation, and getting the order wrong fails the upgrade.
Plugin files are held by the running daemon (a rename probe on the plugin directory returns Access to the path is denied), and on Windows swapping packages can hit EBUSY. So the safe path is "stop DSH → upgrade → start".
On this machine the runtime is launched by a scheduled task running as built-in Administrator with InteractiveToken, so the runtime process is high-integrity: Stop-Process / taskkill from a normal window is denied, and the process's command line is not even readable.
The switch script, however, must run in a normal (non-elevated) window, otherwise the new runtime inherits high integrity and you reproduce the same trap.
The result is a forced order: admin disables the watchdog → admin kills the process → normal window swaps the tree → admin re-enables the watchdog. The watchdog fires every 1 minute, so if the order slips once (it relaunches the old runtime), the upgrade dies at step 1. I hit exactly that.
Plugin management is unavailable precisely where automation is needed. In an agent session under a workspace sandbox, any dsh plugin command fails immediately, because the profile lives outside the workspace and the command needs to write a lock file there:
Error: EPERM … open '<DSH_HOME>\profiles\web\package.json.lock'
[sandbox: file access denied under workspace-write mode]
So install/upgrade/remove can only be done from an elevated window, and "let the agent upgrade its own plugins" — the most natural path — is closed. To upgrade one plugin I wrote a six-step elevation script plus a plugin compatibility health check.
What I think this means [hypothesis]
[hypothesis] The plugin API has no stability tiering and no version-matrix commitment, so plugin authors can only defend themselves passively by declaring peer ranges — and users absorb the entire upgrade cost. Evidence that the need is real: the plugin I upgraded solved this on its own by declaring only two supported lines (0.1.5-rc.1 and 0.1.7-rc.2) and explicitly marking the middle one as transitional.
Suggestions
A compatibility contract and a version matrix. Say which plugin APIs are stable and which are experimental, and which DSH release lines are supported. Formalise what that plugin author did by hand.
Startup compatibility health check with actionable errors. Today I have to grep stderr for skipping profile bundle. The UI should say: this plugin, because of this API change, run this command.
Make plugin crashes readable. A minified React error is zero information for a user; plugin error boundaries should carry the plugin name and version.
Fix the version semantics / provide the upgrade command.dsh plugin upgrade --latest that rewrites the spec and installs in one go, and a UI hint when a plugin is pinned by its spec.
Swap packages without stopping the service (atomic or deferred replacement) so upgrading a plugin doesn't require downtime.
A legitimate plugin-management path for agent sessions (or an explicit authorisation entry point) — otherwise this step can never be automated.
Publish the roadmap and the stability commitment. This is the answer to "what is the main direction of the product": what is core, what is experimental, and what plugin authors may rely on. Speaking as a daily user, that is the single document I most want to read.
3. Structural analysis: the model is overconfident, and negative conclusions are the dangerous ones
What I observed
In a real delivery project the model had to infer data structures from web pages and internal APIs. The errors cluster around one pattern: when reverse-engineering structure from an unstructured interface, the model states a confident but wrong conclusion.
Case 1 — read the wrong field, concluded "it doesn't exist", lost a whole round. The route to a module lives in the menu node's link field. The model read url for the entire round (always empty) and concluded "the menu has no promotion module" — a conclusion that then went into the handover document. The next round, reading the right field, found the entry point (a separate domain, embedded as an iframe by design).
Case 2 — "no menu entry" read as "no permission", almost a permanent wrong conclusion. That account's privilege menu contains no "promotion" entry, so the model concluded "this brand has no promotion data worth looking at". In reality, no menu entry ≠ no API permission — the gateway authorises by login session and returned 7-day, 30-day and same-day figures plus balance. That wrong conclusion also went into the handover document. Without a second round of verification it would have poisoned every later round.
Case 3 — ignoring parameter semantics produced data inflated N times. One endpoint's recentDays=1ignores the beginTime/endTime you pass and always returns the latest day. The model's "fetch day by day and sum" therefore produced figures inflated by a factor of N — and the numbers looked perfectly normal.
Case 4 — a batch of silent correctness bugs, each proven by machine. All of these were caught by assertions or human review, not by the model:
Bug
Effect
Template read storeName/liveStatus; the data actually has name/status
store table showed IDs, status column always empty
#N/A in a source spreadsheet treated as "closed"
closed-store count overstated
City name written both as "杭州" and "杭州市"
filter silently dropped nearly all matching stores (single digits vs several hundred)
Daily chart X axis reversed
trend read backwards
A caption hard-coding one project's figures
both projects displayed the same wrong sentence
String-based replace expanding join("$$") into join("$")
chart silently corrupted
Why I consider this the most expensive item
One wrong attribution = a whole round (half a day) redone. I ended up maintaining 31 gotchas in handover documents as "re-doing this costs 30+ minutes".
The worst part is not the rework, it's the contamination: a wrong conclusion recorded in a handover document is inherited by every later round and reads like verified fact.
The most dangerous class is silent error: numbers and charts look plausible while the definition is wrong — and I use them to make business decisions.
My attribution
Measured: the errors cluster in "infer structure from an unstructured interface", and negative conclusions (absent / doesn't exist / no permission / request was never sent) fail clearly more often than positive ones.
[hypothesis] The model states "I did not find it in this attempt" as "it does not exist": it does not doubt the completeness of its own search, and it does not separate verified fact from unverified assumption in its output.
Suggestions
Require cross-verification for negative conclusions. When the output is of the form "there is no X / it doesn't exist / no permission / the request was never sent", require at least two independent evidence paths (different field name, different entry point, a control account) before it may land as a conclusion — otherwise it must be explicitly labelled "not found ≠ does not exist".
Make structural analysis produce an artifact, not a sentence. Record which fields were read, from which request, what is measured and what is assumed, in a re-runnable file. We already patch this by hand: the project produced a structure-analysis report of 47 measured facts + 11 unverified assumptions + 28 new findings, plus a re-runnable field-contract checker. That format should be built in.
Generate consumer-side assertions from the inferred contract, so "read the wrong field" fails before shipping instead of showing an empty column. Every bug in the table above was found by human review, not by tooling.
Signal repeated work. When the same file/goal is modified repeatedly, or the same endpoint is probed again with contradicting conclusions, surface it: "this area has been attempted N times, results conflict".
What I'm asking for
Confirmation of which of the above behaviours are intended and which are bugs worth filing (especially: IM↔Web session sharing, ^0.x spec semantics, plugin management being unavailable to agent sessions).
A public compatibility contract for plugin authors, and a clear answer on API stability tiers.
A roadmap / stability statement — what is core, what is experimental, what plugin authors can rely on.
I'm happy to turn any of these into a proper issue with a minimal reproduction, or to test a fix on this setup.
Environment for everything above: DSH Web 0.1.7-rc.2 (profile=web), Windows, multi-plugin; session-query and skills findings re-verified today on that exact version.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi all — I've been using DSH Web as my main working environment for about a month, and I want to give back the three problems that cost me the most time. Each one comes with measurements, and I've separated what I measured from what I merely suspect.
TL;DR — (1) Long-session cost didn't disappear, it turned into manual work I now maintain by hand. (2) One patch-level upgrade of DSH took six steps, including elevation and downtime, because plugin compatibility has no contract. (3) In structural-analysis tasks the model is overconfident, and its negative conclusions ("there is no X") are the expensive ones, because nobody re-checks a "no".
A companion post has five smaller gaps with full reproduction steps: [Feedback] Five smaller gaps with reproducible steps.
0. Who I am and what this sample is
profile=web, on Windows; long-lived runtime; multiple third-party plugins; a mix of Web GUI sessions and IM (WeCom) group sessions feeding the same session.1. Long sessions: the cost didn't disappear, it turned into manual work
What I measured
Cost structure of one heavy session:
So cost ≈ average context size × number of calls, and is almost independent of how much is actually produced.
The part I think matters more than the ratio
To work around the context ceiling, I built an external memory system by hand:
The context cost didn't go away — it became the hours I spend maintaining 187 KB of handover documents. And it still isn't enough.
Long-session side effects, same session
tool/call710,tool/result720)Result: the UI visibly degrades, and the user is effectively pushed into starting a new session (losing continuity) or living with the lag.
The session log is not searchable
The session log is appended multi-frame zstd. A one-shot decompress returns only the first frame (the file looks empty); streaming decompression dies on frame two with
Unknown frame descriptor. What I ended up doing is scanning for the magic number28 B5 2F FD, splitting frames and decompressing them one by one: 4.69 MB → 1,946 frames → 16.68 MB → 3,549 JSONL records. Until you reverse that, session history is a black box for both the user and the model.Suggestions
0.1.7-rc.2/profile=web:dsh-session-queryanddsh-session-query-sqlite(FTS5 backend) are installed, but the model-facing packagedsh-tool-session-queryis not. Consequence: in a long session the model has no built-in way to look up "what did the user actually ask earlier" — it either relies on what's still in context, or writes its own script to decompress the session log (which is what I do).dsh session grep <pattern>.2. Plugin compatibility: one patch-level upgrade took six steps
What I measured
Three plugin compatibility failures in a single patch-level upgrade, each printing
dsh: skipping profile bundleon every start — and note this only goes to stderr, never to.out, sogrepin the normal log finds nothing:0.1.6-alpha.2→ the plugin had to be upgraded before DSH would load it.link:plugin whose source had been moved away, leaving a broken symlink innode_modules→ had to be removed from the manifest.bundles.An upstream icon-vocabulary rename broke a plugin, and the error is unreadable. The new harness renamed
ui-primitivesicons as a whole (Icon*Outline16→Icon*OutlineRegular), so named imports in the plugin becameundefinedand the panel crashed on render. All the user sees is:That error tells you nothing about which plugin or which API changed; I had to go read the plugin author's release notes to find out. Worse, the same root cause reports a different number across versions (
#138→#130), which reads like a new problem.Version semantics are counter-intuitive, so upgrades silently do nothing. The manifest said
^0.53.3. Under0.xsemantics that pins the minor line (>=0.53.3 <0.54.0), so0.57.0can never be installed —installreports success and the version does not move. I only found out because the panel kept crashing.Upgrading requires downtime and elevation, and getting the order wrong fails the upgrade.
Access to the path is denied), and on Windows swapping packages can hitEBUSY. So the safe path is "stop DSH → upgrade → start".InteractiveToken, so the runtime process is high-integrity:Stop-Process/taskkillfrom a normal window is denied, and the process's command line is not even readable.Plugin management is unavailable precisely where automation is needed. In an agent session under a workspace sandbox, any
dsh plugincommand fails immediately, because the profile lives outside the workspace and the command needs to write a lock file there:So install/upgrade/remove can only be done from an elevated window, and "let the agent upgrade its own plugins" — the most natural path — is closed. To upgrade one plugin I wrote a six-step elevation script plus a plugin compatibility health check.
What I think this means [hypothesis]
[hypothesis] The plugin API has no stability tiering and no version-matrix commitment, so plugin authors can only defend themselves passively by declaring peer ranges — and users absorb the entire upgrade cost. Evidence that the need is real: the plugin I upgraded solved this on its own by declaring only two supported lines (
0.1.5-rc.1and0.1.7-rc.2) and explicitly marking the middle one as transitional.Suggestions
skipping profile bundle. The UI should say: this plugin, because of this API change, run this command.dsh plugin upgrade --latestthat rewrites the spec and installs in one go, and a UI hint when a plugin is pinned by its spec.3. Structural analysis: the model is overconfident, and negative conclusions are the dangerous ones
What I observed
In a real delivery project the model had to infer data structures from web pages and internal APIs. The errors cluster around one pattern: when reverse-engineering structure from an unstructured interface, the model states a confident but wrong conclusion.
Case 1 — read the wrong field, concluded "it doesn't exist", lost a whole round. The route to a module lives in the menu node's
linkfield. The model readurlfor the entire round (always empty) and concluded "the menu has no promotion module" — a conclusion that then went into the handover document. The next round, reading the right field, found the entry point (a separate domain, embedded as an iframe by design).Case 2 — "no menu entry" read as "no permission", almost a permanent wrong conclusion. That account's privilege menu contains no "promotion" entry, so the model concluded "this brand has no promotion data worth looking at". In reality, no menu entry ≠ no API permission — the gateway authorises by login session and returned 7-day, 30-day and same-day figures plus balance. That wrong conclusion also went into the handover document. Without a second round of verification it would have poisoned every later round.
Case 3 — ignoring parameter semantics produced data inflated N times. One endpoint's
recentDays=1ignores thebeginTime/endTimeyou pass and always returns the latest day. The model's "fetch day by day and sum" therefore produced figures inflated by a factor of N — and the numbers looked perfectly normal.Case 4 — a batch of silent correctness bugs, each proven by machine. All of these were caught by assertions or human review, not by the model:
storeName/liveStatus; the data actually hasname/status#N/Ain a source spreadsheet treated as "closed"join("$$")intojoin("$")Why I consider this the most expensive item
My attribution
Suggestions
What I'm asking for
^0.xspec semantics, plugin management being unavailable to agent sessions).I'm happy to turn any of these into a proper issue with a minimal reproduction, or to test a fix on this setup.
Environment for everything above: DSH Web
0.1.7-rc.2(profile=web), Windows, multi-plugin; session-query and skills findings re-verified today on that exact version.中文版(点击展开)
中文版
TL;DR —— ①长会话的成本没有消失,它变成了我手工维护的工时;②一次补丁级升级要动六步(含提权与停机),因为插件兼容没有契约;③在结构分析类任务上模型过度自信,而否定性结论("没有 X")才是最贵的,因为没人会去验证"没有"。
一、长会话:成本变成了手工劳动。 本机实测一次高强度会话:584 次调用、约 2.74 亿 token、≈¥31.66,其中 98.6% 花在重读上下文,模型输出仅 0.17%。为绕开上下文上限,我自建了一套外置记忆:两份交接文档(605 行 / 66 KB 与 1,229 行 / 111 KB)、31 条"别再踩"的硬坑、一个自写的会话检索工具、一套自写的插件兼容体检。上下文成本没有消失,它变成了我维护 187 KB 交接文档的工时。 同一条会话的其它代价:日志 4.69 MB 压缩 / 16.68 MB 展开、3,549 条记录、界面要渲染 644 个气泡 + 1,430 张工具卡片、Web 进程常驻 599 MB。会话日志是多帧追加 zstd(一次性解压只出第一帧,流式解压报
Unknown frame descriptor),我靠扫魔数28 B5 2F FD切出 1,946 帧才能检索。另外:面型模型的会话检索包缺失(今天在0.1.7-rc.2上复核:dsh-session-query与dsh-session-query-sqlite都在,dsh-tool-session-query不在),所以超长会话里模型没有任何内置手段回查"用户前面说过什么"。建议:上下文预算可见 + 分层归档;把面向模型的会话检索纳入默认组合;公开会话格式并提供dsh session grep;内置轮次折叠/虚拟滚动;把"交接摘要"原生化;明确 IM 与 Web 共用 session id 是有意还是默认。二、插件兼容:一次补丁级升级要动六步。 单次升级暴露 3 个兼容问题,且
dsh: skipping profile bundle只写 stderr、不写.out。上游把图标词表整体改名(Icon*Outline16→Icon*OutlineRegular)就让插件崩,用户只看到Minified React error #138 / #130——不知道是哪个插件、哪条 API 变了;同一根因在不同版本还报不同编号。^0.53.3在0.x下锁的是次版本,装不到0.57.0,install会"成功但版本没动"。插件文件被运行中的守护进程占用,Windows 上换包可能EBUSY,所以更稳的路径是"先停、再升、再起";而本机运行时由计划任务以内置 Administrator + InteractiveToken 拉起,进程是高完整性,普通窗口杀不掉,切换脚本又必须用普通窗口跑 → 正确顺序被逼成「管理员禁看门狗 → 管理员杀进程 → 普通窗口切换 → 管理员恢复看门狗」,看门狗每 1 分钟触发一次,顺序错一次就死在第 1 步。更麻烦的是 agent 会话(工作区沙箱)里任何dsh plugin命令必然失败(profile 在工作区外,写package.json.lock被拒),"让 agent 自己升级插件"这条最自然的路走不通。建议:兼容性契约与版本矩阵(那个插件作者已经自己在声明两条支持线了,应由官方推成规范);启动期兼容体检 + 可执行报错(哪个插件 / 因为哪条 API / 跑什么命令);插件崩溃报错带上插件名与版本;dsh plugin upgrade --latest直接改 spec 并安装;不停机换包;给 agent 会话一条合法的插件管理通道;公开 roadmap 与稳定性承诺——这是"亟需确立产品主要方向"的答案,也是我作为用户最想读到的一份文档。三、结构分析:模型过度自信,否定性结论最危险。 四个案例:①路由在菜单节点的
link字段,模型整轮都在读url(全空),于是判定"菜单里没有该模块"并写进交接文档——下一轮才找到入口;②权限菜单里没有"推广"入口 → 模型判定"该品牌没有推广数据",实测菜单无入口 ≠ API 无权限,网关按登录会话放行并正常返回了数据——这个错误结论也已写进交接文档,若无复核会污染后续所有轮次;③某接口的recentDays=1会忽略传入的起止时间、恒返回最新一天,"逐日取再求和"得到 N 倍虚高且数值看起来正常;④一批静默数据错误(模板读错字段导致列空白、#N/A被当成"未营业"、"杭州/杭州市"两种写法筛掉了绝大多数本应命中的门店、X 轴倒序、写死数值、字符串替换把join("$$")变成join("$"))。代价:单次错误归因 = 一整轮返工;最贵的是污染(错误结论被后续每一轮继承且看起来像已验证事实);危害最大的是静默错误(数值正常但口径错,而我拿它做经营决策)。我的归因(事实):错误集中在"从非结构化界面反推结构",且否定性结论的出错率明显更高;**(假设)**模型把"这次没找到"表达成了"不存在",缺少对自身检索完备性的怀疑,也没有把实测事实与未验证假设分开输出。建议:否定性结论强制交叉验证(两条独立证据路径,否则必须标注"未找到 ≠ 不存在");结构分析产物化(我手工产出的格式是 47 条实测事实 + 11 条待验证假设 + 28 条新发现 + 一个可复跑的字段契约核对脚本,这个格式应当内置);自动生成消费侧断言让错误在上线前报错;识别"重复做功"信号并提示。边界声明:以上是单人单机样本,环境(Windows + 多插件 + 沙箱 + 计划任务)比多数用户更极端,数字是我这种用法的量级而非平均值;标注"假设"的部分尚未验证。
All reactions