Repository navigation
v0.1.38
Release Notes — v0.1.38
Provider usage reports now survive transport format quirks
- What you saw. Some relay channels (GLM, and potentially others served by one-api/new-api-style gateways) report token usage with non-standard JSON types — counts as strings (
"prompt_tokens": "12345"), whole floats, or invalid optional fields (null detail values, fractional costs). hellogrok previously dropped the entire usage measurement on any format defect, forcing Grok Build onto its byte-based estimate. For Chinese and JSON-heavy conversations that estimate overshoots real token counts by 25% or more, so auto-compaction fired far too early — at roughly 50% real usage instead of the configured threshold. The premature summary dropped all tool results, the model re-read files it had already processed, context ballooned again, and a second compaction stalled the session entirely. - What changed. A usage rectifier now repairs provider measurements before they reach Grok Build's token ledger and auto-compaction meter. Stringly-typed counts, whole floats, and
json.Numbervalues are normalized to integers; invalid optional decorations (null detail values, non-map detail containers, fractionalcost_in_usd_ticks) are removed without poisoning the core measurement; unrecoverable core counts (negative, fractional, overflowing, non-numeric) are deleted so the downstream validator drops the measurement instead of trusting corrupt data. Repairs apply across Chat Completions, Responses, and Messages, in streaming and non-streaming responses, and in protocol-translation paths, and are logged asusage rectifiedproxy notes so you can see exactly what was fixed.
Stream rectifier architecture unified
- The three frame-rewriting stream rectifiers (Chat tool calls, Messages tool blocks, Responses function calls) now share a common
streamFrameRectifiercontract, making it straightforward to add rectification for future protocols without touching the dispatch pipeline.
Restart both hellogrok executables after upgrading. If you use a GLM channel, watch the proxy log for usage rectified notes — they confirm the channel's usage reports were previously being dropped and are now reaching Grok Build correctly.
发布说明 — v0.1.38
供应商用量上报现在能经受住传输格式缺陷
- 你看到的现象。 部分中转渠道(GLM,以及可能其他由 one-api/new-api 类网关服务的渠道)上报 token 用量时使用非标准 JSON 类型——计数字段是字符串(
"prompt_tokens": "12345")、整值浮点,或可选字段格式无效(null 详情值、浮点 cost)。hellogrok 此前遇到任何格式缺陷都会丢弃整份用量测量,迫使 Grok Build 回落到字节估算。对中文和 JSON 密集会话,该估算比真实 token 数偏高 25% 以上,导致自动压缩在真实用量仅约 50% 时就提前触发。过早的摘要丢弃了全部工具结果,模型重新读取已处理过的文件,上下文再次膨胀,第二次压缩直接让会话停滞。 - 本次变化。 新增用量整流器,在供应商测量到达 Grok Build 的 token 账本与自动压缩计量表之前进行修复。字符串数字、整值浮点和
json.Number统一归一为整数;无效的可选装饰字段(null 详情值、非对象详情容器、浮点cost_in_usd_ticks)单独剥除而不污染核心测量;不可恢复的核心计数(负数、小数、溢出、非数字)删除后由下游校验按不完整测量丢弃。整流覆盖 Chat Completions、Responses、Messages 三种协议,流式与非流式响应,以及协议互转路径,修复记录以usage rectified代理日志输出,可精确看到修了什么。
流式整流器架构统一
- 三个帧改写型流式整流器(Chat 工具调用、Messages 工具块、Responses 函数调用)现在共享统一的
streamFrameRectifier契约,新增协议整流时无需改动分发管线。
升级后请重启两个 hellogrok 可执行文件。如果你使用 GLM 渠道,留意代理日志中的 usage rectified 记录——它们证实该渠道的用量上报此前被丢弃,现在已能正确到达 Grok Build。