Repository navigation
v0.1.39
Release Notes — v0.1.39
Long relay stalls no longer kill the whole turn
- What you saw. Some relay channels (GLM and others) stream heartbeat bytes for many minutes without producing any content while queued or thinking — stalls of 3+ minutes with heartbeat-only frames have been observed. Grok Build's content-progress timer only counts real content (text, reasoning, tool-call deltas, terminal signals); heartbeats and empty deltas do not reset it. After 600 seconds without content it ends the turn with
IdleTimeout, which is classified non-retryable — the 15-attempt budget never engages and the task dies. - What changed. Two layers of protection:
- Managed idle deadline. Proxied channels without an
inference_idle_timeout_secsvalue now get 1800 seconds materialized temporarily (restored byte-for-byte on stop; first-partyapi.deepseek.comroutes are excluded so their remote metadata stays authoritative), widening the tolerated single stall from 10 to 30 minutes. Explicit per-model or global[models]values remain user-owned and are never rewritten. - Downstream content watchdog. Every streaming path (native Chat/Messages/Responses passthrough and the protocol-translation projections) mirrors Grok Build's per-protocol content classification and fires one 30-second margin ahead of the client's own content timer. On fire the proxy closes the upstream body and emits a retryable
proxy_stream_error, so a genuinely stalled stream re-enters Grok Build's native 15-attempt retry budget instead of hitting the non-retryable client classification. Watchdog activity is logged asSSE stalled: no content reached Grok Build for …. Explicit timeouts below 60 seconds disable the watchdog to respect fast-fail choices.
- Managed idle deadline. Proxied channels without an
Net effect
Together with the existing soft-failure absorb window, a single request now survives: transient pre-header failures (retried inside the proxy), long heartbeat-only stalls (waited out up to ~30 minutes), and a truly stuck stream (converted to a retryable error with up to 15 client retries). Unattended long-running tasks no longer break on relay jitter.
Restart both hellogrok executables after upgrading. If you previously set inference_idle_timeout_secs for a channel, your value is kept as-is; channels without one log inference idle timeout model=… managed=1800s at startup.
发布说明 — v0.1.39
长时间停顿的中转不再杀死整轮任务
- 你看到的现象。 部分中转渠道(GLM 等)在排队或长推理时会持续发心跳字节却长时间不产生任何内容——已观测到超过 3 分钟的纯心跳停顿。Grok Build 的内容进度计时器只认真实内容(文本、推理、工具增量、终态信号),心跳和空增量不算进度;600 秒无内容即以
IdleTimeout结束本轮,而该错误被归类为不可重试——15 次重试预算用不上,任务直接中断。 - 本次变化。 两层防护:
- 托管空闲时限。 未配置
inference_idle_timeout_secs的代理渠道会被临时补为 1800 秒(停止时逐字节恢复;api.deepseek.com官方路由除外,其远端元数据保持权威),单次停顿的等待窗口从 10 分钟放宽到 30 分钟。你显式配置的 per-model 或全局[models]值永远是用户所有,不会被改写。 - 下游无内容看门狗。 所有流式路径(Chat/Messages/Responses 原生透传与协议互转)按 Grok Build 各协议的内容判定口径镜像计时,并始终比客户端自己的内容计时器提前 30 秒触发。触发时代理关闭上游连接并发出可重试的
proxy_stream_error——真正卡死的流因此进入 Grok Build 的原生 15 次重试预算,而不是撞上不可重试的客户端判定。日志表现为SSE stalled: no content reached Grok Build for …。小于 60 秒的显式超时会禁用看门狗,尊重快速失败的选择。
- 托管空闲时限。 未配置
综合效果
配合既有的软故障吸收窗口,单个请求现在可以挺过:响应头前的瞬态故障(代理内重试)、长时间心跳停顿(等待至多约 30 分钟)、以及彻底卡死(转为可重试错误,客户端最多重试 15 轮)。无人值守的长任务不再因中转抖动而中断。
升级后请重启两个 hellogrok 可执行文件。若你此前为渠道手动配置过 inference_idle_timeout_secs,该值保持原样;未配置的渠道启动日志会出现 inference idle timeout model=… managed=1800s。