Repository navigation
Releases: hellowind777/hellogrok
Release list
v0.1.41
Release Notes — v0.1.41
Self-healing path for unnamed tool calls
- Unnamed Chat tool calls no longer die as a silent terminal
NotFound. When a streamed call keeps an empty name after shape inference (for example an empty name plus arguments missing the opening{with trailing garbage), hellogrok now routes it to a deterministic carrier instead of handing Grok Build an empty name the model never sees fed back. - Parseable arguments go to
run_terminal_command(commandcarries the raw text); corrupt arguments go toread_file(target_filecarries the raw text). Both carriers fail cleanly, so the raw arguments return through atool_resultand the model regenerates the call on the next turn. The route is logged asunresolved-name-routed(name=…). When neither carrier is declared on the request, the call is left untouched.
Upgrade the running proxy to pick up the new behavior; the next affected turn recovers without a new session.
发布说明 — v0.1.41
无名工具调用进入自愈路径
- 无名 Chat 工具调用不再以空名称死于终端态
NotFound。 当流式调用在形状推断后仍无名称(例如空名称叠加缺开头{"且带尾部垃圾的参数),hellogrok 会将其路由到确定性载体,而不是把模型永远看不到回执的空名称交给 Grok Build。 - 可解析参数走
run_terminal_command(command携带原文),损坏参数走read_file(target_file携带原文)。 两个载体都会干净失败,原文经tool_result回喂模型,下一轮由模型重新生成调用。路由记录为unresolved-name-routed(name=…)。当请求未声明这两个载体时保持原样透传。
升级运行中的代理即可生效;下一次命中的回合会自动恢复,无需新建会话。
v0.1.40
Release Notes — v0.1.40
Shorter managed stall window, uniform across channels
- Managed
inference_idle_timeout_secslowered from 1800 to 900 seconds. The 30-minute ceiling traded too much failure-detection latency for stall tolerance. Fifteen minutes still covers twice DeepSeek's documented ten-minute queue and multiples of the observed three-minute relay stalls, and the downstream watchdog (deadline minus 30 seconds) already converts a true stall into a retryableproxy_stream_error, so the extra wait mostly delayed discovery. - The managed timeout now applies to every proxied channel, including first-party
api.deepseek.comroutes. It is a resilience projection, not channel metadata, so it is no longer special-cased. Explicit per-model or global[models]values remain user-owned and win on every channel alike. The proxy's DeepSeek-specific 660-second upstream byte window is unaffected.
Restart both hellogrok executables after upgrading; the rewritten channel values take effect at the next proxy start. Channels you configured explicitly keep their value unchanged.
发布说明 — v0.1.40
托管停顿窗口缩短,且对所有渠道统一生效
- 托管
inference_idle_timeout_secs从 1800 秒降为 900 秒。 30 分钟的上限为停顿容忍付出的失败感知延迟过高。15 分钟仍是 DeepSeek 文档排队上限(10 分钟)的 1.5 倍,也数倍于实测的 3 分钟中转停顿;而下游看门狗(时限减 30 秒)本来就会把真正的卡死转为可重试的proxy_stream_error,多出的等待主要只是推迟了发现。 - 托管超时现在对所有代理渠道统一生效,包括
api.deepseek.com官方路由。 它是韧性投影而非渠道元数据,因此不再被特判。显式的 per-model 或全局[models]值在所有渠道上同样保持用户所有且优先。代理对 DeepSeek 官方的 660 秒上游字节窗口不受影响。
升级后请重启两个 hellogrok 可执行文件;渠道改写值在下次代理启动时生效。你显式配置过的渠道保持原值不变。
v0.1.39
Release Notes — v0.1.39
Long relay stalls no longer kill the whole turn
- What you saw. Some relay channels (GLM and others) stream heartbeat bytes for many minutes without producing any content while queued or thinking — stalls of 3+ minutes with heartbeat-only frames have been observed. Grok Build's content-progress timer only counts real content (text, reasoning, tool-call deltas, terminal signals); heartbeats and empty deltas do not reset it. After 600 seconds without content it ends the turn with
IdleTimeout, which is classified non-retryable — the 15-attempt budget never engages and the task dies. - What changed. Two layers of protection:
- Managed idle deadline. Proxied channels without an
inference_idle_timeout_secsvalue now get 1800 seconds materialized temporarily (restored byte-for-byte on stop; first-partyapi.deepseek.comroutes are excluded so their remote metadata stays authoritative), widening the tolerated single stall from 10 to 30 minutes. Explicit per-model or global[models]values remain user-owned and are never rewritten. - Downstream content watchdog. Every streaming path (native Chat/Messages/Responses passthrough and the protocol-translation projections) mirrors Grok Build's per-protocol content classification and fires one 30-second margin ahead of the client's own content timer. On fire the proxy closes the upstream body and emits a retryable
proxy_stream_error, so a genuinely stalled stream re-enters Grok Build's native 15-attempt retry budget instead of hitting the non-retryable client classification. Watchdog activity is logged asSSE stalled: no content reached Grok Build for …. Explicit timeouts below 60 seconds disable the watchdog to respect fast-fail choices.
- Managed idle deadline. Proxied channels without an
Net effect
Together with the existing soft-failure absorb window, a single request now survives: transient pre-header failures (retried inside the proxy), long heartbeat-only stalls (waited out up to ~30 minutes), and a truly stuck stream (converted to a retryable error with up to 15 client retries). Unattended long-running tasks no longer break on relay jitter.
Restart both hellogrok executables after upgrading. If you previously set inference_idle_timeout_secs for a channel, your value is kept as-is; channels without one log inference idle timeout model=… managed=1800s at startup.
发布说明 — v0.1.39
长时间停顿的中转不再杀死整轮任务
- 你看到的现象。 部分中转渠道(GLM 等)在排队或长推理时会持续发心跳字节却长时间不产生任何内容——已观测到超过 3 分钟的纯心跳停顿。Grok Build 的内容进度计时器只认真实内容(文本、推理、工具增量、终态信号),心跳和空增量不算进度;600 秒无内容即以
IdleTimeout结束本轮,而该错误被归类为不可重试——15 次重试预算用不上,任务直接中断。 - 本次变化。 两层防护:
- 托管空闲时限。 未配置
inference_idle_timeout_secs的代理渠道会被临时补为 1800 秒(停止时逐字节恢复;api.deepseek.com官方路由除外,其远端元数据保持权威),单次停顿的等待窗口从 10 分钟放宽到 30 分钟。你显式配置的 per-model 或全局[models]值永远是用户所有,不会被改写。 - 下游无内容看门狗。 所有流式路径(Chat/Messages/Responses 原生透传与协议互转)按 Grok Build 各协议的内容判定口径镜像计时,并始终比客户端自己的内容计时器提前 30 秒触发。触发时代理关闭上游连接并发出可重试的
proxy_stream_error——真正卡死的流因此进入 Grok Build 的原生 15 次重试预算,而不是撞上不可重试的客户端判定。日志表现为SSE stalled: no content reached Grok Build for …。小于 60 秒的显式超时会禁用看门狗,尊重快速失败的选择。
- 托管空闲时限。 未配置
综合效果
配合既有的软故障吸收窗口,单个请求现在可以挺过:响应头前的瞬态故障(代理内重试)、长时间心跳停顿(等待至多约 30 分钟)、以及彻底卡死(转为可重试错误,客户端最多重试 15 轮)。无人值守的长任务不再因中转抖动而中断。
升级后请重启两个 hellogrok 可执行文件。若你此前为渠道手动配置过 inference_idle_timeout_secs,该值保持原样;未配置的渠道启动日志会出现 inference idle timeout model=… managed=1800s。
v0.1.38
Release Notes — v0.1.38
Provider usage reports now survive transport format quirks
- What you saw. Some relay channels (GLM, and potentially others served by one-api/new-api-style gateways) report token usage with non-standard JSON types — counts as strings (
"prompt_tokens": "12345"), whole floats, or invalid optional fields (null detail values, fractional costs). hellogrok previously dropped the entire usage measurement on any format defect, forcing Grok Build onto its byte-based estimate. For Chinese and JSON-heavy conversations that estimate overshoots real token counts by 25% or more, so auto-compaction fired far too early — at roughly 50% real usage instead of the configured threshold. The premature summary dropped all tool results, the model re-read files it had already processed, context ballooned again, and a second compaction stalled the session entirely. - What changed. A usage rectifier now repairs provider measurements before they reach Grok Build's token ledger and auto-compaction meter. Stringly-typed counts, whole floats, and
json.Numbervalues are normalized to integers; invalid optional decorations (null detail values, non-map detail containers, fractionalcost_in_usd_ticks) are removed without poisoning the core measurement; unrecoverable core counts (negative, fractional, overflowing, non-numeric) are deleted so the downstream validator drops the measurement instead of trusting corrupt data. Repairs apply across Chat Completions, Responses, and Messages, in streaming and non-streaming responses, and in protocol-translation paths, and are logged asusage rectifiedproxy notes so you can see exactly what was fixed.
Stream rectifier architecture unified
- The three frame-rewriting stream rectifiers (Chat tool calls, Messages tool blocks, Responses function calls) now share a common
streamFrameRectifiercontract, making it straightforward to add rectification for future protocols without touching the dispatch pipeline.
Restart both hellogrok executables after upgrading. If you use a GLM channel, watch the proxy log for usage rectified notes — they confirm the channel's usage reports were previously being dropped and are now reaching Grok Build correctly.
发布说明 — v0.1.38
供应商用量上报现在能经受住传输格式缺陷
- 你看到的现象。 部分中转渠道(GLM,以及可能其他由 one-api/new-api 类网关服务的渠道)上报 token 用量时使用非标准 JSON 类型——计数字段是字符串(
"prompt_tokens": "12345")、整值浮点,或可选字段格式无效(null 详情值、浮点 cost)。hellogrok 此前遇到任何格式缺陷都会丢弃整份用量测量,迫使 Grok Build 回落到字节估算。对中文和 JSON 密集会话,该估算比真实 token 数偏高 25% 以上,导致自动压缩在真实用量仅约 50% 时就提前触发。过早的摘要丢弃了全部工具结果,模型重新读取已处理过的文件,上下文再次膨胀,第二次压缩直接让会话停滞。 - 本次变化。 新增用量整流器,在供应商测量到达 Grok Build 的 token 账本与自动压缩计量表之前进行修复。字符串数字、整值浮点和
json.Number统一归一为整数;无效的可选装饰字段(null 详情值、非对象详情容器、浮点cost_in_usd_ticks)单独剥除而不污染核心测量;不可恢复的核心计数(负数、小数、溢出、非数字)删除后由下游校验按不完整测量丢弃。整流覆盖 Chat Completions、Responses、Messages 三种协议,流式与非流式响应,以及协议互转路径,修复记录以usage rectified代理日志输出,可精确看到修了什么。
流式整流器架构统一
- 三个帧改写型流式整流器(Chat 工具调用、Messages 工具块、Responses 函数调用)现在共享统一的
streamFrameRectifier契约,新增协议整流时无需改动分发管线。
升级后请重启两个 hellogrok 可执行文件。如果你使用 GLM 渠道,留意代理日志中的 usage rectified 记录——它们证实该渠道的用量上报此前被丢弃,现在已能正确到达 Grok Build。
v0.1.37
Release Notes — v0.1.37
Reasoning that discusses think tags no longer leaks into the visible reply
- What you saw. v0.1.36 stripped inline think spans wherever they appeared in the stream, but it treated every literal tag occurrence as a delimiter — including tags a model wrote as quoted prose. Reasoning that discussed the tags themselves (for example analyzing a thinking-leak defect) terminated its own span at the first quoted
</think>, and the rest of that reasoning leaked into the visible reply. Quoted tags inside visible text were stripped as well, so replies citing the tags arrived mangled. A related streaming-path defect: the non-streaming content peeler classified unbalanced closing-tag tails and removed stray closing tags on every SSE delta, consuming a closing tag before the turn-level state machine ever saw it, which could flash split tags in the reply. - What changed. A tag occurrence wrapped in backticks or double quotes is now treated as prose about the tags, not as a stream delimiter: quoted tags never open or close a span, only unquoted stray closing tags are removed, and quoted tags inside visible text reach the client intact. The unbalanced-tail classification and stray-close removal in the content path now apply only to complete message objects (those carrying a role); streaming deltas leave classification to the turn-level state machine as designed.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.37
谈论 think 标签的推理不再泄漏进可见正文
- 你看到的现象。 v0.1.36 开始剥离流中任意位置的 think span,但它把每个字面标签出现都当作分隔符——包括模型作为引用正文写出的标签。当推理在讨论标签本身(例如分析思考泄漏缺陷)时,第一个被引用的
</think>会提前终止自己的 span,其余推理随之泄漏进可见回复。可见正文中被引用的标签同样被剥掉,引用标签的回复到达时已不完整。另有一个流式路径的相关缺陷:非流式内容剥离器在每个 SSE 增量上执行无平衡闭标签推理尾判定与游离闭标签删除,会在回合级状态机看到之前消耗掉闭标签,可能让分裂的标签在回复中闪现。 - 本次变化。 被反引号或双引号包住的标签出现现在按"讨论标签的正文"处理,不再作为流分隔符:被引用的开/闭标签不会打开或关闭 span,只有未被引用的游离闭标签会被删除,可见正文中被引用的标签原样送达。内容路径上的无平衡推理尾判定与游离闭标签删除只作用于完整消息对象(携带 role 的对象);流式增量按设计把分类交还给回合级状态机。
升级后请重启两个 hellogrok 可执行文件。
v0.1.36
Release Notes — v0.1.36
Text-only channels no longer end a turn on vision_not_supported
- What you saw. When a conversation carried image content — for example a
read_fileof a picture, whose result Grok Build renders as image blocks — every later request forwarded those image parts to the upstream. A text-only model then rejected the whole request with400 vision_not_supported("The request model is not multimodal … does not support image input"), and the turn died on the red error. - What changed. hellogrok now recognizes that rejection (
400with avisionerror code, or anot multimodal/does not support imagemessage) on any of the three upstream protocols. It replaces every image content part (image_url,input_image,image, including images inside tool results) with a one-line text placeholder that tells the model the visual payload was omitted, and retries the request once. The turn continues without the images instead of failing. Channels are deliberately not remembered as text-only, so the decision is re-derived from each upstream rejection and a channel whose upstream later gains vision support keeps working unchanged.
Third-party tool calls with wrong or missing arguments are normalized or dropped before dispatch
- What you saw. Third-party models occasionally emit Grok Build parameter names from their own training prior —
read_filecalled withtarget_pathinstead oftarget_file, for instance — and Grok Build's strict schema validation answeredFailed to parse arguments for tool …: missing field …, marked the call failed in the TUI, and burned a round trip. Rarer relay/model glitches emit a tool call whose argument fragments never arrive at all; the empty arguments object then fails the same validation with a guaranteedmissing fielderror. - What changed.
- Parameter aliases are normalized against the tools declared on the request:
target_pathnow maps totarget_file,target_directory, orfile_path(whichever the declared schema actually has), joining the existing alias table. Rewrites only apply to properties the tool declares, so unrelated JSON is never touched. - A streamed call that accumulates no argument fragments at all is discarded before Grok Build dispatch when the declared tool requires properties — an empty object can never satisfy it — and the defect stays visible as an
empty-args-discarded(name=…)proxy log note. Calls to tools without required properties keep their empty arguments, which is the valid zero-argument convention.
- Parameter aliases are normalized against the tools declared on the request:
Inline reasoning no longer leaks into the visible reply
- What you saw. Thinking models emit several chain-of-thought phases per turn, and some relays misroute a reasoning tail into
contentwith only the closing tag. The previous think-tag stripper only handled a<think>…</think>block at the very head of the answer, so second-phase spans, unbalanced</think>tails, and stray or delta-split closing tags appeared verbatim in the reply bubble. - What changed. The Chat think-tag state machine now strips spans wherever they appear in the stream: text before an opening tag is emitted immediately, the span is buffered until its closing tag and routed to the reasoning channel, an unbalanced closing tag at the head of a turn classifies the text before it as reasoning, and stray or delta-split closing tags are removed or held instead of being shown. The non-streaming path peels every span from the answer content. Reasoning that arrives after visible reply text is still dropped by design, because Grok Build renders reasoning only as a prefix Thought.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.36
纯文本渠道不再因 vision_not_supported 终结回合
- 你看到的现象。 会话一旦携带图像内容——例如
read_file读取图片、Grok Build 将结果渲染为图像块——之后的每个请求都会把这些图像部分转发给上游。纯文本模型随即以400 vision_not_supported("The request model is not multimodal … does not support image input")拒绝整个请求,回合停在红色报错上。 - 本次变化。 hellogrok 现在能在三种上游协议上识别该拒绝(
400且 code 含vision,或消息含not multimodal/does not support image):把每个图像内容部分(image_url、input_image、image,含 tool result 内的图像)替换为一行文本占位符(告知模型视觉载荷被省略),并重试一次请求。回合不再失败,而是不带图像继续。渠道不会被记为纯文本——判定每次都来自上游的真实拒绝,上游日后获得视觉能力时渠道无需任何改动即可恢复透传。
参数名错误或零参数的第三方工具调用在分发前被归一或丢弃
- 你看到的现象。 第三方模型偶尔按自身训练先验发出 Grok Build 的参数名——例如
read_file写成target_path而非target_file——Grok Build 的严格 schema 校验回以Failed to parse arguments for tool …: missing field …,TUI 将该调用标为失败并消耗一轮往返。更罕见的中继/模型 glitch 会发出一个参数片段完全未到达的工具调用;空参数对象在同一校验下必然报missing field。 - 本次变化。
- 参数别名按该请求已声明的工具归一:
target_path现在映射到target_file、target_directory或file_path(取已声明 schema 实际拥有的那个),并入既有别名表。改写只作用于工具已声明的属性,无关 JSON 不受影响。 - 一个参数片段都未累积的流式调用,在已声明工具含必填属性时于 Grok Build 分发前丢弃——空对象永远无法满足必填字段——缺陷以
empty-args-discarded(name=…)代理日志注记保持可见。无必填属性的工具保留空参数,那是合法的零参调用约定。
- 参数别名按该请求已声明的工具归一:
行内推理不再泄漏进可见回复
- 你看到的现象。 思考模型一个回合产出多段 CoT,部分中继还会把推理尾只带闭标签地误路由进
content。旧的 think 标签剥离器只处理答案最开头的<think>…</think>块,于是第二阶段 span、无开标签的</think>推理尾、游离或跨增量分裂的闭标签都原样出现在回复气泡里。 - 本次变化。 Chat think 标签状态机现在剥离流中任意位置的 span:开标签前的文本立即下发,span 缓冲到闭标签后归入推理通道;回合开头无开标签的闭标签把其前的文本判定为推理;游离或跨增量分裂的闭标签被删除或挂起而不是显示。非流式路径剥离答案内容中的每一个 span。可见回复之后到达的推理仍按设计丢弃,因为 Grok Build 只把推理渲染为前缀 Thought。
升级后请重启两个 hellogrok 可执行文件。
v0.1.35
Release Notes — v0.1.35
Tool calls whose first stream delta was lost are repaired instead of failing
- What you saw. Some third-party relays intermittently omit a tool call's first streamed delta — the frame carrying
function.nameand the opening{"of the arguments. Grok Build then dispatched an empty tool name with head-truncated arguments and reportedAgent tried calling a tool that doesn't exist, feeding a parse error back into the model and burning a retry round (observed repeatedly on GLM-family channels through relays). - What changed. hellogrok now repairs that damage in its shared tool-compatibility layer, for all three upstream protocols (
chat_completions,messages,responses), streaming and non-streaming alike:- Arguments truncated at the head are restored when prefixing
{"(or{) yields exactly one complete JSON object. - Empty names are inferred from the argument key set against the tools declared on that request; the shared
command+descriptionshape, which matches bothrun_terminal_commandandmonitor, resolves deterministically torun_terminal_command. Anything still ambiguous is left unrepaired rather than guessed. - Persisted history replayed in later requests (including
/resume) receives the same repair, and the repaired name is backfilled onto the matching tool-result message, so broken records no longer reach the upstream verbatim.
- Arguments truncated at the head are restored when prefixing
Out-of-order and name-less tool blocks are held until they can be resolved
- Chat Completions. Tool frames are now emitted at the stream terminal (
[DONE], stream end, or error frame) instead of atfinish_reason. Grok Build's chat accumulator reads the whole stream regardless of chunk order, so argument fragments a relay sends afterfinish_reasonare merged into complete arguments instead of being silently dropped with tail-truncated parameters. Held inline reasoning still flushes atfinish_reason, so visible latency of thought and text is unchanged. - Messages.
tool_useblocks are held betweencontent_block_startandcontent_block_stop; the accumulated input JSON resolves the name (and repairs its prefix) before the block is re-emitted as start + one completeinput_json_delta+ stop. Unclosed blocks flush atmessage_stop/error. - Responses.
function_callitems are held betweenresponse.output_item.addedandresponse.output_item.done; the done frame's complete item additionally serves as a second source for name and arguments, so a truncated argument stream loses to a complete terminal item. Unclosed items flush atresponse.completed/incomplete/failed.
Diagnostics
- Tool-call deltas that arrive after an early flush (error terminals) and are therefore discarded are logged per call index as
late-tool-deltas-discarded(index=N), making a relay that emitsfinish_reasonbefore its last argument fragments visible in the proxy log instead of surfacing only as a Grok Build parse failure.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.35
丢失首个流式增量的工具调用现在会被修复,而不是报错
- 你看到的现象。 部分第三方中继会间歇性丢掉工具调用的第一个流式增量——携带
function.name与参数开头{"的那一帧。Grok Build 随后以空工具名和缺开头的参数进行分发,报出Agent tried calling a tool that doesn't exist,并把解析错误回喂给模型、消耗一轮重试(在经中继的 GLM 系渠道上反复出现)。 - 本次变化。 hellogrok 在共享的工具兼容层修复这类损坏,覆盖三种上游协议(
chat_completions、messages、responses)及流式与非流式:- 参数缺开头时,补
{"(或{)后若能解析为恰好一个完整 JSON 对象,则恢复前缀。 - 空名称按该请求已声明工具的参数键集合推断;
command+description这一同时匹配run_terminal_command与monitor的共享形状确定性地解析为run_terminal_command。仍无法判定时保持不修复,不做猜测。 - 后续请求回放的已持久化历史(含
/resume)接受同样的修复,修复后的名称回填到对应 tool result 消息,损坏记录不再原样送达上游。
- 参数缺开头时,补
乱序与缺名工具块扣留到可解析为止
- Chat Completions。 工具帧改在流终止(
[DONE]、流末或错误帧)发出,而非finish_reason时刻。Grok Build 的 chat 累加器读取整条流、不依赖 chunk 顺序,因此中继在finish_reason之后补发的参数片段会并入完整参数,而不是被静默丢弃留下尾部截断。持有中的行内推理仍在finish_reason时发出,思考与正文的可见延迟不变。 - Messages。
tool_use块在content_block_start与content_block_stop之间扣留;累积的 input JSON 先解析名称(并修复前缀),再以"开始帧 + 单条完整input_json_delta+ stop"重发。未闭合块在message_stop/error时冲刷。 - Responses。
function_call项在response.output_item.added与response.output_item.done之间扣留;done 帧的完整 item 同时作为名称与参数的第二来源,截断的流上参数让位于完整的终止 item。未闭合项在response.completed/incomplete/failed时冲刷。
诊断
- 提前 flush(错误终止)之后到达、因而被丢弃的工具调用增量,按调用索引记录为
late-tool-deltas-discarded(index=N)。中继在最后一个参数片段之前发出finish_reason的缺陷由此在代理日志中可见,而不是只表现为 Grok Build 的解析失败。
升级后请重启两个 hellogrok 可执行文件。
v0.1.34
Release Notes — v0.1.34
Retryable Responses response.failed events are now absorbed instead of streamed
- What you saw. A relay under rate-limit or concurrency pressure returned a Responses SSE whose only terminal was a retryable
response.failed(rate_limit_exceeded,Concurrency limit exceeded, overload). hellogrok streamed that failure to Grok Build immediately, so the turn failed even though a retry seconds later would have succeeded — and the retry then opened a second stream for the same request. - What changed. While nothing has been written to the client, such a failed-only SSE is now withheld inside the absorb window (response headers and early frames stay buffered, up to 32 frames) and replayed with the same exponential backoff and
Retry-Afterhandling as HTTP soft failures. Grok Build keeps its full retry budget; only an exhausted window passes the failure through, still retryable.absorb_retry_max_secs = 0streams the failed event immediately. Deterministicresponse.failedevents (authentication, invalid request/model) still stream through without retry. - Retry classification widened. Concurrency-limit rejections (
concurrency,concurrency limit, plus the corresponding Chinese phrases) now classify as transient, and error envelopes nested underresponse.errorare recognized the same as top-levelerrorobjects.
Client disconnects no longer become stream errors
- What you saw. Closing a Grok Build session mid-stream could leave a
proxy_stream_errorin the log or on a late reader, suggesting an upstream failure that never happened. - What changed. Messages, Chat Completions, native, and Responses streams now distinguish a client abort (failed client write, canceled request context) from an upstream failure: aborts are logged as
aborted by clientand emit no stream error. Truly truncated upstream streams still emitproxy_stream_error.
Strict errors for stream-shape mismatches
- An SSE response to a non-streaming request, or a streaming response to Grok Build's fixed non-streaming WebSearchClient request, now returns a non-retryable
502naming the mismatch instead of forwarding an undecodable body.
Quieter, safer diagnostics
- Upstream HTTP errors and Responses
response.failed/errorevents are logged as structured summaries (type,code,message) with bearer tokens, key assignments, andsk-values redacted and long messages truncated. - Benign upstream-model mismatches log once per channel/protocol/configured/upstream pair; conflicts and invalid declarations always log.
- When Grok Build sends an official catalog name such as
grok-4.6on a non-xAI custom channel, the proxy logs a warning with body size, tool count, and session presence to aid/resumediagnosis. First-partyapi.x.airoutes and custom IDs such asgrok4.6-sevnxare excluded.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.34
可重试的 Responses response.failed 事件现在会被吸收,而不是直接透传
- 你看到的现象。 中转在限流或并发压力下返回的 Responses SSE,其唯一终态是可重试的
response.failed(rate_limit_exceeded、Concurrency limit exceeded、过载)。此前 hellogrok 会立即把该失败透传给 Grok Build,本轮直接失败——而几秒后重试本可成功;重试还会为同一请求再开一条流。 - 本次变更。 在尚未向客户端写入任何内容时,这类“只有失败终态”的 SSE 现在进入吸收窗口:响应头和早期帧先缓冲(最多 32 帧),并按与 HTTP 软故障相同的指数退避与
Retry-After规则在代理内重放。Grok Build 的重试预算分毫不动;只有窗口耗尽后才以可重试形式透传。absorb_retry_max_secs = 0则直接透传该失败事件。确定性response.failed(鉴权、无效请求或模型)仍直接透传,不重试。 - 重试分类放宽。 并发限制拒绝(
concurrency、concurrency limit及对应中文“并发限制/并发超限”)现在归为瞬态故障;嵌套在response.error下的错误信封与顶层error同等识别。
客户端断开不再被记为流错误
- 你看到的现象。 在流式传输中途关闭 Grok Build 会话,日志或迟到的读取者可能看到
proxy_stream_error,像是上游出了故障,而实际只是客户端已离开。 - 本次变更。 Messages、Chat Completions、原生与 Responses 流现在区分客户端中止(客户端写入失败、请求上下文取消)与上游故障:中止只记录为
aborted by client,不产生流错误。上游真正截断的流仍会产生proxy_stream_error。
流形态不匹配现在明确报错
- 非流式请求收到 SSE 响应,或 Grok Build 固定的非流式 WebSearchClient 请求收到流式响应时,返回不可重试的
502并说明不匹配原因,不再转发无法解码的正文。
更安静、更安全的诊断日志
- 上游 HTTP 错误与 Responses
response.failed/error事件按结构化摘要(type、code、message)记录;bearer 令牌、key 赋值和sk-值会被脱敏,超长消息会被截断。 - 良性上游模型不一致按渠道/协议/配置模型/上游模型组合只记录一次;冲突与无效声明每次都记录。
- 当 Grok Build 在非 xAI 自定义渠道上发送
grok-4.6这类官方目录名时,代理会记录一条警告(含请求体大小、工具数量和会话状态),便于诊断/resume选路。官方api.x.ai路由与grok4.6-sevnx这类自定义 ID 不在警告范围内。
升级后请重启两个 hellogrok 可执行文件。
v0.1.33
Release Notes — v0.1.33
Completed relay streams are no longer failed for omitting a wire trailer
- Not tied to one channel or model. The repair is per protocol, not per provider name. It applies to every custom channel: Chat Completions, Messages, and Responses, whatever model ID the relay uses.
- What you saw. Grok Build showed
Server error | Retrying (attempt N)and thenServer error: Something went wrong on our side. Wait a minute and send again., while the relay dashboard stayed HTTP 200. The yellow banner hid the real cause: aproxy_stream_errorthat the stream had ended without a terminal event. - What was actually happening. Many relays return a full SSE stream (Chat
finish_reason, Messagesstop_reason, Responses output items or a completedstatus) then close without the trailer Grok Build's decoder wants ([DONE],message_stop,response.completed). Grok Build can complete Chat/Messages on a clean close; Responses cannot. hellogrok previously injectedproxy_stream_errorin every case, which Grok Build retries up to 15 times — each retry a new billed 200 — then surfaces the generic server-error copy. The retry countdown is backoff, not a first-token timeout. - The fix. On a clean close, hellogrok synthesizes the missing trailer for that protocol instead of rewriting the stream as an error. Truly truncated streams (no stop signal and no output), idle timeouts, and read failures still emit
proxy_stream_error.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.33
已完成的中转流不再因省略协议尾帧而被判失败
- 不绑某个渠道或模型。 按协议处理,不看供应商名字。所有自定义渠道都适用:Chat Completions、Messages、Responses,中转侧用什么模型 ID 都一样。
- 你看到的现象。 Grok Build 显示
Server error | Retrying (attempt N),随后变成Server error: Something went wrong on our side. Wait a minute and send again.,而中转后台一直是 HTTP 200。黄字横幅把真实原因藏掉了:流「没有终态事件」的proxy_stream_error。 - 实际发生了什么。 很多中转会返回完整 SSE 流(Chat 的
finish_reason、Messages 的stop_reason、Responses 的 output item 或已 completed 的status),然后不发 Grok Build 解码器要的尾帧([DONE]、message_stop、response.completed)就关连接。Chat/Messages 在干净关闭时 Grok 自己能收尾;Responses 不能。此前 hellogrok 一律注入proxy_stream_error,Grok Build 最多重试 15 次——每次都在中转侧再记一笔 200——预算耗尽后打出通用 Server error。重试倒计时是退避,不是首字超时。 - 修复。 干净关闭时按协议补发缺失尾帧,而不是把流改写成错误。真正的截断(没有任何停止信号也没有任何输出)、空闲超时和读失败仍发
proxy_stream_error。
升级后请重启两个 hellogrok 可执行文件。
v0.1.32
Release Notes — v0.1.32
DeepSeek models no longer get an implicit backend-search default after DeepSeek retired hosted web search
-
The old default now routes
web_searchinto a dead end. hellogrok used to project models on the exact first-partyapi.deepseek.comendpoint assupports_backend_search = truewhen the field was omitted, based on DeepSeek's documented provider-hosted web search. DeepSeek has since retired that capability: the current Responses API documentsweb_searchand every other built-in tool type as ignored, and its response format no longer includesweb_search_call. With the old default, aweb_searchcall went to the DeepSeek channel, the endpoint silently answered without executing any search, and Grok Build showed a completed turn with no sources — while the client-search fallback chain ([models].web_search,GROK_WEB_SEARCH_MODEL, the authenticated official fallback) was bypassed entirely because the channel looked search-capable. -
An omitted field now behaves exactly like every other provider. The implicit default is removed. Omission preserves Grok Build catalog behavior uniformly: client
web_searchresolves from[models].web_search,GROK_WEB_SEARCH_MODEL, or the authenticated official fallback, and if none is available the model reports that web search is unavailable instead of pretending to search. hellogrok no longer writessupports_backend_search = trueinto the active configuration for an omitted DeepSeek channel. Explicit values keep their meaning:trueis honored as a routing declaration (useful on relayed endpoints that still implement a real search extension; the first-party endpoint silently ignores it), andfalsestays opted out unless that route is selected as the default search model. -
No other DeepSeek behavior changed. The
[1m]Anthropic Messages alias, thinking-mode normalization, the 660-second idle policy covering the documented ten-minute queue, Bearer/X-Api-Keyauthentication, and the Chat-to-Responses search bridge dialect are untouched. Configuration that explicitly setsupports_backend_searchbefore upgrading behaves identically.
Also reflected in documentation: DeepSeek's 2026-09-10 announcement replaced the V4 Flash generation with V4.1 Flash under the model ID deepseek-flash (retired deepseek-v4-flash / deepseek-v4-flash-vision-exp names still route to it), and V4 Pro service continues. Existing DeepSeek channel configurations need no changes.
Restart both hellogrok executables after upgrading.
发布说明 — v0.1.32
DeepSeek 下线 hosted 搜索后,其模型不再获得隐式后端搜索默认值
-
旧默认值现在会把
web_search引向死路。 此前 hellogrok 会对精确指向官方api.deepseek.com端点且未配置该字段的模型,临时投影为supports_backend_search = true——依据是 DeepSeek 文档承诺的供应商托管搜索。DeepSeek 已下线该能力:当前 Responses API 将web_search及全部内建工具类型标记为忽略,响应格式也不再包含web_search_call。在旧默认值下,web_search调用被路由到 DeepSeek 渠道,端点不执行任何搜索就静默作答,Grok Build 显示一个没有任何来源的"已完成"回合——而客户端搜索回退链([models].web_search、GROK_WEB_SEARCH_MODEL、已登录官方账号回退)被整体绕过,因为该渠道看起来具备搜索能力。 -
字段缺省时的行为现在与其他供应商完全一致。 隐式默认值已移除。缺省一律保留 Grok Build 模型目录行为:客户端
web_search依次从[models].web_search、GROK_WEB_SEARCH_MODEL或已登录官方账号回退中解析;全都不可用时,模型如实报告无法使用 web 搜索,而不是假装搜索。hellogrok 不再为缺省该字段的 DeepSeek 渠道向活动配置写入supports_backend_search = true。显式值的含义不变:true仍作为路由声明被尊重(对仍实现真实搜索扩展的中转端点有用;官方端点会静默忽略),false仍保持关闭,除非该渠道被选为默认搜索模型。 -
其他 DeepSeek 行为均未改变。
[1m]Anthropic Messages 别名、思考模式规范化、覆盖官方十分钟排队的 660 秒空闲策略、Bearer/X-Api-Key鉴权,以及 Chat 到 Responses 的搜索桥接方言全部保持原样。升级前已显式设置supports_backend_search的配置行为不变。
文档同步反映:DeepSeek 2026-09-10 公告以模型 ID deepseek-flash 上线 V4.1 Flash 一代(已退役的 deepseek-v4-flash / deepseek-v4-flash-vision-exp 名称仍路由到该模型),V4 Pro 服务继续提供。既有 DeepSeek 渠道配置无需任何改动。
升级后请重启两个 hellogrok 可执行文件。