Repository navigation
v0.1.29
Release Notes — v0.1.29
Upstream 408s and short relay outages no longer kill unattended turns
-
Upstream
408is treated as what it really is. Between two servers there is no "client was too slow" — a relay or Cloudflare edge emits 408 for an edge-side timeout while the origin is queued or restarting, semantically a 504. Grok Build classifies a literal 408 as terminal and fails the turn on first sight. hellogrok now retries a 408 inside the absorb window and remaps an exhausted one to a retryable504, so the client keeps its native retry budget instead of dying. This was the exact failure that interrupted long-running agent turns on busy relays. -
Soft-failure absorption is always on, deterministic 4xx is configurable. Retryable
5xx,429, 408 edge-timeout pages, Cloudflare origin-TLS525/526, and response-header timeouts are absorbed for every channel, because those clear on their own and the client's retry of them only burns time. Deterministic rejections (authentication, permission, invalid request/model) pass through immediately with the provider's explanation by default. A new global[models]error_resilience = "balanced"setting additionally absorbs those 4xx failures with a fixed 30s-then-60s wait sequence, so an unattended turn survives relay-side token rotation, permission fixes, config reloads, and brief deploys. A reasoning-history rejection is never absorbed at any setting: it reports a foreign conversation state the user must see. -
A long-hidden failure says how stale it is. An error that outlasted a long absorb window is the last attempt's snapshot, not a live one. hellogrok now stamps
X-Hellogrok-Absorb-Delayon the passthrough response and prefixes structured error messages with a one-line[hellogrok: upstream stayed failing for ...]note, so the delay is visible instead of masquerading as a fresh failure. Once the window is exhausted, Grok Build's own retry indicator shows the error reason (<headline> | Retrying (N/M)) while its native budget continues — the failure never hides behind a bare retrying state.
Relay TLS failures now reach the dead-channel breaker
The opt-in breaker (dead_channel_fail_fast = true) only counts dial-level failures, but its TLS classification covered only certificate-verification errors. It now counts every handshake-failure form Go actually produces — remote/local alert errors, tls.RecordHeaderError, and the HTTPS-to-plain-HTTP scheme mismatch — so a channel with broken TLS termination fast-fails instead of burning the full client retry budget.
Responses thought-gate hardening
- Post-answer reasoning can no longer leak through the
output_item.doneframe. The event now honors the same drop/strip rules as itsaddedcounterpart;reasoning_summary_part.added/donetext goes through the protocol self-talk filter; a post-answer<think>-only content part is dropped. - Relay thinking variants are gated, not leaked. Items and content parts typed
thinking,reasoning_summary, orredacted_thinking— what relays that bridge Anthropic-style blocks through protocol conversion emit — pass through the same sanitize/drop rules as officialreasoning, in both streamed frames and the terminalresponse.completedframe. - Fail-loud on the unknown. The gate is now table-driven with one registered rule per event type; an unregistered event that visibly carries reasoning is intercepted and logged instead of leaking a second Thought under the reply. Unrelated unknown events keep passing untouched.
Configuration-recovery fixes
- Invalid-TOML recovery no longer deletes a user-edited root
subagents.enableddotted line, and drops management of feature-flag lines whose entire[features]section the user deleted — both now match the parse path. A new invariant test pins the two restore paths to the same decision for the same input. - Requests caught mid-flight by a proxy stop receive the structured
503 proxy_stoppeddiagnostic instead of a one-off retryable502, so stale sessions know to reselect a model.
Every change stays inside the existing safety contract: repairs happen before Grok Build's strict validators see the wire, deterministic refusals remain non-retryable so a rejected request never re-enters a retry loop, and error_resilience is the only new configuration — optional, global, off by default.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.29
上游 408 与中转短时故障不再中断无人值守的任务
-
上游
408按其真实语义处理。 服务器之间不存在"客户端太慢"——中转或 Cloudflare 边缘在源站排队、重启时发出的是边缘超时页,语义上等价 504。Grok Build 把字面 408 归类为终局错误并在首次出现即中断本轮。hellogrok 现在会在吸收窗口内重试 408,窗口耗尽后改写为可重试的504透传,客户端保留原生重试预算而不再直接失败。这正是此前中断长时间智能体任务的那种故障。 -
软故障吸收始终生效,确定性 4xx 可配置。 可重试
5xx、429、408 边缘超时页、Cloudflare 源站 TLS525/526和响应头超时对每个渠道都吸收——它们会自行恢复,交给客户端重试只是空耗时间。确定性拒绝(鉴权、权限、无效请求/模型)默认立即透传供应商解释。新增全局[models]error_resilience = "balanced"设置可额外把这些 4xx 也纳入吸收窗口,按固定 30 秒后 60 秒的节奏等待,无人值守的任务因此能挺过中转侧令牌轮换、权限修复、配置热加载和短暂部署。推理历史被拒绝在任何设置下都不吸收:它报告的是用户必须看到的异源会话状态。 -
被隐藏较久的失败会标明其时效。 挺过长时间吸收窗口的错误是最后一次尝试的快照而非实时状态。hellogrok 现在在透传响应上标记
X-Hellogrok-Absorb-Delay,并在结构化错误消息前加上[hellogrok: upstream stayed failing for ...]一行提示,延迟可见而不再伪装成刚发生的失败。窗口耗尽后,Grok Build 的重试指示会显示错误原因(<原因> | Retrying (N/M))并继续原生重试——失败永远不会藏在光秃秃的"重试中"状态后面。
中转 TLS 失败现在能触发死渠道熔断
可选熔断器(dead_channel_fail_fast = true)只统计拨号级失败,但其 TLS 分类此前只覆盖证书校验错误。现在它统计 Go 实际产生的全部握手失败形态——远端/本地 alert 错误、tls.RecordHeaderError、HTTPS 指向明文 HTTP 的协议错配——TLS 终止配置错误的渠道会快速失败,而不是烧光客户端全部重试预算。
Responses 思维门控加固
- 正文后的思维不再能从
output_item.done帧泄漏。 该事件现在遵循与其added对应帧相同的丢弃/剥离规则;reasoning_summary_part.added/done文本经过协议自语过滤;正文后纯<think>内容块被丢弃。 - 中转思维变体被门控而非泄漏。 类型为
thinking、reasoning_summary、redacted_thinking的 item 与内容块——中转把 Anthropic 风格块经协议转换桥接时的产物——在流式帧与终帧response.completed中都按官方reasoning的清洗/丢弃规则处理。 - 对未知事件响亮失败。 门控改为表驱动、每种事件类型一条注册规则;未注册但明显携带思维的事件被拦截并记录日志,而不是在用户回复下方泄漏第二个 Thought。无关的未知事件照旧透传。
配置恢复修复
- 非法 TOML 恢复不再删除用户编辑过的根表
subagents.enableddotted 行;用户整体删除[features]节时也不再对其中的功能开关行行使管理权——两者现在都与解析路径一致。新增不变量测试把两条恢复路径钉在同一决策上。 - 被代理停止时正好在途的请求收到结构化的
503 proxy_stopped诊断,而不是一次性的可重试502,过期会话因此知道要重新选择模型。
全部变更都维持在既有安全契约内:修复发生在 Grok Build 的严格校验器看到线路数据之前,确定性拒绝保持不可重试(被拒绝的请求不会重进重试循环),error_resilience 是唯一新增配置——可选、全局、默认关闭。
升级后请重启两个 hellogrok 可执行文件。