[Proposal] Provider-neutral prompt token pressure + local overflow -> compact-retry #8374
Replies: 6 comments
|
中文摘要 问题: 官方 DSH 用固定字符/token 启发式估算上下文,中文长会话和工具结果常被低估。本地 llama.cpp 在真正超窗时返回 exceed_context_size_error,但 stock DSH 常把它当成普通 400,已有的 compact-retry 路径接不住。 社区方案(不改官方 core): 独立 dsh-plugin,通过 llm/stream 钩子实现: 请求前调用本地服务 /tokenize 精算 token,超阈值主动触发压缩 实测: 官方桌面版 DSH + 64K llama.cpp(Bonsai),token 约 5.6 万超阈值 → 自动压缩 → 回落约 1.8 万,会话可继续。 希望上游考虑: 可选 countPromptTokens 协议(provider 中立) |
Clarification / update (plugin defaults)The linked reference plugin has been updated so defaults are provider-neutral:
Repo: https://github.com/tianyiming1/dsh-plugin-local-prompt-bridge ( No change to the proposed core direction ( |
|
Clarification: plugin defaults are now provider-neutral (routes: [], global overflow rewrite). Bonsai/18200 was only the verification machine, not a required default. See v0.1.1 on the reference repo. |
Update (plugin v0.1.2): soft vs hard tokenize pressureFollow-up from real usage on official DSH + local 64K llama.cpp (Bonsai2): Bug we hit: synthesizing
Plugin fix (v0.1.2):
Repo: https://github.com/tianyiming1/dsh-plugin-local-prompt-bridge (commits Please treat soft vs hard pressure as part of any future core |
Update v0.1.3Soft/hard refined further:
Repo: https://github.com/tianyiming1/dsh-plugin-local-prompt-bridge |
Update v0.1.4 — align with Cursor/Codex ~90% compactAfter comparing production agents (Cursor ~90% auto-summarize, Codex Plugin change:
This is token-budget management (industry), not a client “remaining KV %” gauge. Repo: https://github.com/tianyiming1/dsh-plugin-local-prompt-bridge |
Uh oh!
There was an error while loading. Please reload this page.
Problem
Stock
dsh-token-meterprices prompts with a fixed ~4 characters/token heuristic. Dense-script (CJK) sessions and large tool schemas are systematically undercounted. Local OpenAI-compat servers (e.g. llama.cpp) then reject the real request with:Two gaps follow:
agent/request-error+CONTEXT_WINDOW_EXCEEDED→ compact → retry), but bare HTTP 400 +exceed_context_size*is often classified asINVALID_REQUEST, so the recovery path never runs.Reproduced on local 64K llama.cpp routes with Chinese-heavy agent sessions (heuristic far below window, server reports prompt over
n_ctx).Proposed core direction (decoupled)
Optional adapter capability, no engine names in core:
token-meter/ compaction call through the runtime when present; on miss/timeout → keep heuristic.promptTokenCount: true | { endpoint, input, templateEndpoint }.isContextWindowExceededError/ stream classification soexceed_context_size*maps toCONTEXT_WINDOW_EXCEEDEDbefore bare-400 →INVALID_REQUEST.This stays provider-neutral: DeepSeek cloud, gateways, and local servers can each implement the method; core never hard-codes
/tokenizeor llama.cpp.What the community can use today
While external PRs are closed, a stock-safe plugin implements the operational half without forking
@deepseek-ai/*:llm/streamrewrite of overflow-like finish errors →CONTEXT_WINDOW_EXCEEDEDagent/request-errorfallback compact+retry for misclassified overflowsbaseURL; when overthresholdRatio × contextWindow, synthesizeCONTEXT_WINDOW_EXCEEDEDfor stock compact-retryInstall shape: npm /
github:/link:bundle withdsh.bundle.patch, topicdsh-plugin.Reference implementation:
https://github.com/tianyiming1/dsh-plugin-local-prompt-bridge
dsh-pluginpnpm add github:tianyiming1/dsh-plugin-local-prompt-bridge, then add@local/dsh-plugin-local-prompt-bridgeto profilebundlesVerified on official DSH desktop path + 64K llama.cpp (Bonsai): tokenize max above threshold → synthesized
CONTEXT_WINDOW_EXCEEDED→ stock compact-retry recovered session (~56k → ~18k).Ask
countPromptTokensseam acceptable for a future core change?exceed_context_size*land even sooner (small, high leverage)?Happy to split into: (a) overflow classifier only, (b) protocol + meter, (c) reference adapter opt-in — whenever external contributions reopen.
All reactions