Replies: 1 comment
|
这个问题和 SandBase Harness 的 MCP 生命周期设计有一个直接交集:建议把“连接状态”和“工具注册表”分开管理,连接暂时不可用时不要永久丢弃稳定的工具名与 schema。 SandBase Harness 当前的 MCP Manager 采用两层处理:
实现:
对应测试覆盖了“连接断开后自动重连并重试工具调用”以及“永久故障按退避重连后返回失败”。这不能替代你提议的 exhaustion slow-retry:对于跨越整个 fast budget 的长时间故障,仍建议保留工具标识、进入低频后台重试,并在成功 initialize/tools-list 后恢复可用状态;同时要让 close/HMR/会话 teardown 取消慢重试定时器,避免 crash-loop 变成无限 spawn。 |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
@deepseek-ai/dsh-mcp-clientreconnects a dropped MCP server with exponential backoff (0.5s doubling to a 30s cap,maxAttemptsdefault 10). Once that budget is exhausted,scheduleReconnect()disposes all of the server's registered tools and gives up permanently — the log readsgiving up after 10 consecutive failed reconnect attempts — tools unregistered; reload the plugin or restart the Host to reconnect. With the defaults, the cumulative backoff window is only ~2.5 minutes (0.5+1+2+4+8+16+4x30s, computed from the documented parameters), so one ordinary server outage longer than that — a deploy, a restart, a network blip — permanently disables the server's tools for the rest of a long session, even after the server comes back healthy.Background & Evidence
scheduleReconnect()(dsh-mcp-client/lib/index.js, ~L583 in 0.1.1-rc.2) disposes every entry in the disposers map and logs the error quoted above.Failure Mode or Reproduction
Observed: the server's tools remain unregistered for the rest of the session; the log shows the
giving up ... tools unregisterederror. Only a plugin reload or a Host restart brings them back.Proposed Fix
Keep the fast exponential-backoff budget exactly as-is (it still bounds the crash-loop case), but replace permanent unregistration at exhaustion with a slow-retry keepalive:
failedAttempts > maxAttempts, keep the tools registered and re-scheduleconnectGeneration(false)on a fixed slow interval (we use 30 minutes; anything in the 5-30 minute range seems reasonable), logging one warn line per attempt.maxDelayMs, resetfailedAttemptsso the next outage enjoys a fresh fast budget.reconnect.exhaustion = "unregister" | "slow-retry"(default slow-retry), so operators who prefer fail-fast can opt back in.We have a working local patch against 0.1.1-rc.2 implementing the slow-retry approach (a
SLOW_RETRY_DELAY_MSconstant, the rewritten exhaustion branch, and a docstring update) — happy to share the diff in this thread if that would help.Additional Context
@deepseek-ai/dsh0.1.1-rc.2 (developer preview)All reactions