Skip to content

v1.17.0 — Quality scale 1-7, glm-5.2, score recalibration

Choose a tag to compare

@stackbilt-admin stackbilt-admin released this 17 Jun 12:29
· 13 commits to main since this release

What's new

Extended quality scale (1–5 → 1–7)

The previous 1–5 scale had claude-opus-4.6 and llama-3.3-70b both at quality:5, leaving no headroom for near-frontier MoE models like gpt-oss-120b and kimi-k2.6. The new scale:

Score Tier Examples
7 True frontier claude-opus-4.6, claude-opus-4.8
6 Near-frontier MoE gpt-oss-120b, kimi-k2.6/k2.7, deepseek-v4-pro, glm-5.2, zai-glm-4.7
5 Strong large/MoE llama-3.3-70b, qwen3-30b-a3b, glm-4.7-flash
4 Capable mid llama-4-scout, gpt-oss-20b, haiku-class
3 Capable small mistral-small-24b, llama 8B
2 Limited small llama 3B
1 Toy <2B params

Routing effect

gpt-oss-120b (quality:6) now correctly outscores llama-3.3-70b (quality:5) for HIGH_PERFORMANCE and TOOL_CALLING tasks where quality weight is dominant.

New model: @cf/zai-org/glm-5.2

Z.ai flagship agentic coding model. Active on Workers AI. 262K context, direct-response (not CoT), function calling. $1.40/$4.40 per 1M tokens.

Bug fix

@cf/openai/gpt-oss-120b description corrected — this model is OpenAI's own open-weight 117B/5.1B-active MoE (≈ o4-mini tier), not Phi-4 as the previous description implied.

Also catalogued (retired/EOL)

  • @cf/moonshotai/kimi-k2.5 — EOL 2026-05-30, excluded from routing
  • @cf/google/gemma-3-12b-it — EOL 2026-05-30, excluded from routing