Skip to content

v0.0.16

Choose a tag to compare

@github-actions github-actions released this 09 Sep 09:21
· 1128 commits to master since this release

v0.0.16 — 修复 Anthropic 流式计费 usage 双计与缓存语义

修复 Anthropic / Vertex Claude / AWS Bedrock Claude 渠道流式请求中 usage 被双计(input/cache 令牌按约两倍扣费)的问题,并统一 Anthropic 缓存计费语义:cache read 折入 prompt 计价、补齐 AWS 渠道此前完全缺失的 cache 统计。

中文

🐛 问题修复

  • 修复 Anthropic 流式计费「usage 双计」(issue #13):
    • 新版 Messages API 下 message_start 与 message_delta 都携带整段请求的累计 usage(message_delta 重复 message_start 的 input/cache 字段并带上最终 output_tokens)。
    • 原实现按 += 当作增量累加,导致 input_tokens / cache_read_input_tokens 被计两遍,典型场景多扣约 50%(如 3.5-sonnet 下 4522 vs 正确 3000)。
    • 现改为对累计值取 max 合并,并兼容旧形态(message_delta 仅带 output、input 为 0)不丢 message_start 的 input 计数。
  • 修复 cache read 语义错配造成的欠费:Anthropic 的 input_tokens 与 cache_read/cache_creation_input_tokens 互斥且不含彼此,而共享计费公式按 OpenAI「cached ⊆ prompt」扣减(input×(prompt−cached)),导致 cache read 被从 prompt 中重复扣减、按 read 价计费的同时又丢了一次 input 价;现归一化把 read 折入 prompt,公式自动还原出 input×input + readPrice×read 的正确金额。
  • 修复 AWS Bedrock Claude 渠道 cache 完全不参与计费:流式与非流式路径此前都没有把 cache_read_input_tokens 写入 PromptTokensDetails.CachedTokens,现统一补齐。
  • 修复 cache_creation(写入令牌)解析后从不计费:cache_creation_input_tokens 此前仅被解析、未进入 usage;现折入 prompt 按输入价计费(真实写入价约为 1.25×input,此为已知近似,见 #13)。

🔧 重构

  • 新增 relay/adaptor/anthropic/usage.go:ClaudeUsage2OpenAI(Claude→OpenAI usage 归一化)与 MergeClaudeUsage(累计流式事件取 max 合并)。
  • native Anthropic / Vertex AI Claude / AWS Bedrock Claude 三条消费路径统一复用上述 helper,消除三处手写累计逻辑的漂移。

🧪 测试

  • relay/adaptor/anthropic/main_test.go:message_delta fixture 从旧形态(input_tokens: 0)更新为真实累计形态,避免继续掩盖双计 bug。
  • 新增 TestClaudeUsage2OpenAI / TestMergeClaudeUsage 行为单测:覆盖累计序列不双计、旧形态兼容、read/creation 折入、cached ⊆ prompt 不变式。

⚠️ 升级注意事项

  • 零数据库迁移、零配置变更。
  • 计费口径调整:修复后 Anthropic 渠道不再双计,扣费恢复正常;prompt_tokens 消费日志将包含折入的 cache read/creation,口径与 OpenAI 一致(历史消费记录不追溯修正)。
  • 运行验证:go build ./...;go test ./relay/adaptor/anthropic/ ./relay/billing/ratio/ ./model/ ./controller/ ./middleware/。

English

🐛 Bug Fixes

  • Fixed double-counted streaming usage for Anthropic (issue #13):
    • Under the current Messages API, both message_start and message_delta carry cumulative usage for the whole request (message_delta repeats message_start's input/cache fields and adds the final output_tokens).
    • The old code accumulated with += as if each event were incremental, so input_tokens / cache_read_input_tokens were counted twice — typically ~50% over-billing (e.g. 4522 vs the correct 3000 on claude-3.5-sonnet).
    • Events are now merged by max of the cumulative values, and the legacy shape (delta with only output_tokens, zeroed input) still keeps message_start's input count.
  • Fixed cache-read semantic mismatch causing under-billing: Claude's input_tokens is disjoint from cache_read/cache_creation_input_tokens, yet the shared billing formula assumes OpenAI's "cached ⊆ prompt" (input×(prompt−cached)). Cache reads were therefore subtracted twice — charged at read price while also dropping one input-price charge. Reads are now folded into PromptTokens, so the formula yields the correct input×input + readPrice×read.
  • Fixed AWS Bedrock Claude channels never billing cache: neither the streaming nor the non-streaming path wrote cache_read_input_tokens into PromptTokensDetails.CachedTokens; both now do.
  • Fixed cache-creation tokens being parsed but never billed: cache_creation_input_tokens was unmarshalled yet dropped from usage; it is now folded into PromptTokens and billed at input price (actual write price is ≈1.25×input — a documented approximation, see #13).

🔧 Refactor

  • New relay/adaptor/anthropic/usage.go: ClaudeUsage2OpenAI (Claude→OpenAI usage normalization) and MergeClaudeUsage (max-merge of cumulative stream events).
  • The three consumer paths — native Anthropic, Vertex AI Claude, and AWS Bedrock Claude — now share these helpers, removing three divergent hand-written accumulators.

🧪 Tests

  • relay/adaptor/anthropic/main_test.go: the message_delta fixture was updated from the legacy zeroed-input_tokens shape to the real cumulative shape so the double-count bug can no longer hide.
  • Added TestClaudeUsage2OpenAI / TestMergeClaudeUsage behavior tests covering cumulative sequences, legacy-shape compatibility, read/creation folding, and the cached ⊆ prompt invariant.

⚠️ Upgrade Notes

  • Zero database migration and zero configuration changes.
  • Billing behavior: Anthropic channels no longer double count and charge correctly; prompt_tokens in usage logs now includes folded cache read/creation, consistent with OpenAI semantics (historical logs are not retroactively adjusted).
  • Verification: go build ./...; go test ./relay/adaptor/anthropic/ ./relay/billing/ratio/ ./model/ ./controller/ ./middleware/.