Skip to content

v0.1.29

Choose a tag to compare

@github-actions github-actions released this 13 Sep 07:22
· 16 commits to main since this release

Release Notes — v0.1.29

Upstream 408s and short relay outages no longer kill unattended turns

  • Upstream 408 is treated as what it really is. Between two servers there is no "client was too slow" — a relay or Cloudflare edge emits 408 for an edge-side timeout while the origin is queued or restarting, semantically a 504. Grok Build classifies a literal 408 as terminal and fails the turn on first sight. hellogrok now retries a 408 inside the absorb window and remaps an exhausted one to a retryable 504, so the client keeps its native retry budget instead of dying. This was the exact failure that interrupted long-running agent turns on busy relays.

  • Soft-failure absorption is always on, deterministic 4xx is configurable. Retryable 5xx, 429, 408 edge-timeout pages, Cloudflare origin-TLS 525/526, and response-header timeouts are absorbed for every channel, because those clear on their own and the client's retry of them only burns time. Deterministic rejections (authentication, permission, invalid request/model) pass through immediately with the provider's explanation by default. A new global [models] error_resilience = "balanced" setting additionally absorbs those 4xx failures with a fixed 30s-then-60s wait sequence, so an unattended turn survives relay-side token rotation, permission fixes, config reloads, and brief deploys. A reasoning-history rejection is never absorbed at any setting: it reports a foreign conversation state the user must see.

  • A long-hidden failure says how stale it is. An error that outlasted a long absorb window is the last attempt's snapshot, not a live one. hellogrok now stamps X-Hellogrok-Absorb-Delay on the passthrough response and prefixes structured error messages with a one-line [hellogrok: upstream stayed failing for ...] note, so the delay is visible instead of masquerading as a fresh failure. Once the window is exhausted, Grok Build's own retry indicator shows the error reason (<headline> | Retrying (N/M)) while its native budget continues — the failure never hides behind a bare retrying state.

Relay TLS failures now reach the dead-channel breaker

The opt-in breaker (dead_channel_fail_fast = true) only counts dial-level failures, but its TLS classification covered only certificate-verification errors. It now counts every handshake-failure form Go actually produces — remote/local alert errors, tls.RecordHeaderError, and the HTTPS-to-plain-HTTP scheme mismatch — so a channel with broken TLS termination fast-fails instead of burning the full client retry budget.

Responses thought-gate hardening

  • Post-answer reasoning can no longer leak through the output_item.done frame. The event now honors the same drop/strip rules as its added counterpart; reasoning_summary_part.added/done text goes through the protocol self-talk filter; a post-answer <think>-only content part is dropped.
  • Relay thinking variants are gated, not leaked. Items and content parts typed thinking, reasoning_summary, or redacted_thinking — what relays that bridge Anthropic-style blocks through protocol conversion emit — pass through the same sanitize/drop rules as official reasoning, in both streamed frames and the terminal response.completed frame.
  • Fail-loud on the unknown. The gate is now table-driven with one registered rule per event type; an unregistered event that visibly carries reasoning is intercepted and logged instead of leaking a second Thought under the reply. Unrelated unknown events keep passing untouched.

Configuration-recovery fixes

  • Invalid-TOML recovery no longer deletes a user-edited root subagents.enabled dotted line, and drops management of feature-flag lines whose entire [features] section the user deleted — both now match the parse path. A new invariant test pins the two restore paths to the same decision for the same input.
  • Requests caught mid-flight by a proxy stop receive the structured 503 proxy_stopped diagnostic instead of a one-off retryable 502, so stale sessions know to reselect a model.

Every change stays inside the existing safety contract: repairs happen before Grok Build's strict validators see the wire, deterministic refusals remain non-retryable so a rejected request never re-enters a retry loop, and error_resilience is the only new configuration — optional, global, off by default.

Restart both hellogrok executables after upgrading.


发布说明 — v0.1.29

上游 408 与中转短时故障不再中断无人值守的任务

  • 上游 408 按其真实语义处理。 服务器之间不存在"客户端太慢"——中转或 Cloudflare 边缘在源站排队、重启时发出的是边缘超时页,语义上等价 504。Grok Build 把字面 408 归类为终局错误并在首次出现即中断本轮。hellogrok 现在会在吸收窗口内重试 408,窗口耗尽后改写为可重试的 504 透传,客户端保留原生重试预算而不再直接失败。这正是此前中断长时间智能体任务的那种故障。

  • 软故障吸收始终生效,确定性 4xx 可配置。 可重试 5xx、429、408 边缘超时页、Cloudflare 源站 TLS 525/526 和响应头超时对每个渠道都吸收——它们会自行恢复,交给客户端重试只是空耗时间。确定性拒绝(鉴权、权限、无效请求/模型)默认立即透传供应商解释。新增全局 [models] error_resilience = "balanced" 设置可额外把这些 4xx 也纳入吸收窗口,按固定 30 秒后 60 秒的节奏等待,无人值守的任务因此能挺过中转侧令牌轮换、权限修复、配置热加载和短暂部署。推理历史被拒绝在任何设置下都不吸收:它报告的是用户必须看到的异源会话状态。

  • 被隐藏较久的失败会标明其时效。 挺过长时间吸收窗口的错误是最后一次尝试的快照而非实时状态。hellogrok 现在在透传响应上标记 X-Hellogrok-Absorb-Delay,并在结构化错误消息前加上 [hellogrok: upstream stayed failing for ...] 一行提示,延迟可见而不再伪装成刚发生的失败。窗口耗尽后,Grok Build 的重试指示会显示错误原因(<原因> | Retrying (N/M))并继续原生重试——失败永远不会藏在光秃秃的"重试中"状态后面。

中转 TLS 失败现在能触发死渠道熔断

可选熔断器(dead_channel_fail_fast = true)只统计拨号级失败,但其 TLS 分类此前只覆盖证书校验错误。现在它统计 Go 实际产生的全部握手失败形态——远端/本地 alert 错误、tls.RecordHeaderError、HTTPS 指向明文 HTTP 的协议错配——TLS 终止配置错误的渠道会快速失败,而不是烧光客户端全部重试预算。

Responses 思维门控加固

  • 正文后的思维不再能从 output_item.done 帧泄漏。 该事件现在遵循与其 added 对应帧相同的丢弃/剥离规则;reasoning_summary_part.added/done 文本经过协议自语过滤;正文后纯 <think> 内容块被丢弃。
  • 中转思维变体被门控而非泄漏。 类型为 thinking、reasoning_summary、redacted_thinking 的 item 与内容块——中转把 Anthropic 风格块经协议转换桥接时的产物——在流式帧与终帧 response.completed 中都按官方 reasoning 的清洗/丢弃规则处理。
  • 对未知事件响亮失败。 门控改为表驱动、每种事件类型一条注册规则;未注册但明显携带思维的事件被拦截并记录日志,而不是在用户回复下方泄漏第二个 Thought。无关的未知事件照旧透传。

配置恢复修复

  • 非法 TOML 恢复不再删除用户编辑过的根表 subagents.enabled dotted 行;用户整体删除 [features] 节时也不再对其中的功能开关行行使管理权——两者现在都与解析路径一致。新增不变量测试把两条恢复路径钉在同一决策上。
  • 被代理停止时正好在途的请求收到结构化的 503 proxy_stopped 诊断,而不是一次性的可重试 502,过期会话因此知道要重新选择模型。

全部变更都维持在既有安全契约内:修复发生在 Grok Build 的严格校验器看到线路数据之前,确定性拒绝保持不可重试(被拒绝的请求不会重进重试循环),error_resilience 是唯一新增配置——可选、全局、默认关闭。

升级后请重启两个 hellogrok 可执行文件。