Skip to content

Release 0.2.0

Choose a tag to compare

@aarora79 aarora79 released this 11 Sep 20:47
· 20 commits to main since this release
d8b12d5

Release 0.2.0 - Codex joins the harness list, MiniCPM5-2B joins the frontier

September 2026

Version bump: MINOR -- adds codex as a fifth supported coding harness, MiniCPM5-2B to the omp /swe3 frontier, and two benchmarked models that stay off it: Qwen3.6-35B-A3B FP8 on one L40S and Qwen3-Coder-30B-A3B on Bedrock. All of it is additive: no published score or cost from 0.1.0 moved.


Upgrading from 0.1.0

This repo is cloned and run, not deployed, so upgrading is a pull plus a dependency sync.

git pull origin main
git checkout 0.2.0

# Sync the uv project(s) you run. No dependency changed in this release,
# so a sync is only needed if you are coming from a fresh clone:
cd benchmarks && uv sync && cd ..
cd self-hosted/vllm && uv sync && cd ../..

Methodology changes

None. Numbers remain comparable with 0.1.0. Every entry in the three frontier files was compared between the 0.1.0 tag and this release: pareto-frontier-omp-swe3.json gained minicpm5-2b, pareto-frontier-cc-swe3.json and pareto-frontier-combined-swe3.json gained kimi-k3, and not one existing model's score or cost per task changed.

The token-accounting fix in #184 reads as a methodology change but is not one: it only affects the codex agent, which had no published results before this release. The fleet-wide accounting fix that did move numbers shipped in 0.1.0.

Config changes

File Change Action needed
benchmarks/scripts/runner_config.py codex added to VALID_AGENTS, with an is_codex property None. agent still defaults to claude.
benchmarks/config/runner.example.yaml The agent comment now names all five harnesses and how each one's cost is accounted None.
self-hosted/vllm/pricing.json Unchanged None.
benchmarks/scripts/bedrock_pricing.py Qwen3-Coder-30B-A3B priced at $0.1545 / $0.6180 per 1M input / output; cost_usd returns None instead of zero when a caller reports cached tokens against a model that cannot cache None. Adds a rate, changes none.

New models

MiniCPM5-2B on omp / mcp-gateway-registry-v2

A 2.52B dense model scores 42.59 at $0.3993 per task, from pareto-frontier-omp-swe3.json. It completed 19 of 21 tasks; derive-repo-url-from-skill-md and registration-admission-control-gate produced no artifacts and are excluded from the mean rather than averaged in as zeros. Per tier it scores 43.76 trivial, 50.96 low, 40.36 medium, 33.45 high (models.json), so it holds up better on small changes than on architectural ones.

The serving guide is minicpm5-2b.md. Two things matter for anyone serving it: the weights are 5.03 GB in BF16 with a 128K native context, so it fits one L40S with room to spare, and it needs vLLM's minicpm5 tool parser because the model emits XML tool calls rather than JSON. The HuggingFace card still points at SGLang for tool calling, which is out of date.

Its throughput sweep (performance-summary.json) ran on g6e.4xlarge at $1.298/hr and peaked at 82.49 output tokens/sec at concurrency 7, with the cheapest blended rate of $0.04 per 1M tokens at concurrency 5.

Qwen3.6-35B-A3B FP8 on a single L40S, held off the frontier

FP8 on one L40S scores 59.23 under omp and 61.45 under codex over the same 21 tasks, against the BF16 four-GPU run's 59.24. Both runs live under do-not-include/qwen3.6-35b-fp8/ and reach no chart, because the model's only throughput sweep is on g6e.4xlarge rather than the canonical p5en.48xlarge arm the rest of the fleet is priced on, and its token counts are confounded with a 262144 context window against the BF16 run's 200000. The do-not-include README records both runs, their cost per task on the g6e basis ($0.5555 omp, $0.2645 codex at concurrency 10), and the token-accounting convention each harness uses.

The serving guide is qwen3.6-35b-a3b-fp8.md. The useful result: the model runs at its full 256K context on a $1.298/hr single-GPU box at the same quality as its BF16 parent.

Qwen3-Coder-30B-A3B on codex / Bedrock, off the frontier

The first benchmarked model on the codex harness against a hosted API rather than a local server. It scores 42.22 over 19 of 21 tasks at $0.6204 per task, $13.17 for the run, from its run-summary.json. Two model failures are excluded from the mean: cli-custom-egress-oauth-provider-flags filed no artifacts after 187 turns and 3.5M input tokens, and registration-admission-control-gate is missing review.md.

Reaching it needs a translating proxy, and that is a property of Bedrock, not of the harness. codex 0.153.4 speaks only the Responses API, and Bedrock serves Responses for the openai.* family alone: a Qwen request is rejected with The model 'qwen.qwen3-coder-30b-a3b-v1:0' does not support the '/openai/v1/responses' API. Of the two surfaces Bedrock does serve for Qwen, Converse is the one whose usage block can carry cache tokens, so the route is codex to a LiteLLM proxy to Converse. The committed litellm-mantle.yaml registers Qwen as openai/<id> against bedrock-mantle and forwards Responses untouched, so it cannot serve this path.

Every one of its 83,488,690 input tokens was billed fresh. cache_read is 0 on all 21 tasks, because Bedrock refuses a Converse cachePoint for the Qwen family with AccessDeniedException: You invoked an unsupported model or your request did not allow prompt caching, while the same request against claude-haiku-4-5 caches 5,204 tokens. That is what separates $13.17 metered here from the $5.56 a comparable codex run costs against self-hosted Qwen3.6-35B-FP8, where vLLM served 96.4% of the prompt from its prefix cache.

The run reaches no chart on two counts: the harness is codex, which no published frontier scans, and a metered Bedrock bill does not compare as raw dollars with the hardware-derived self-hosted points.


New functionality

codex as a fifth coding harness

--agent codex drives a benchmark run with OpenAI Codex (codex exec --json). Codex has no --skill flag, so the harness inlines the SKILL.md into the prompt, the way it already does for kiro-cli. It works on provider=bedrock through AWS_REGION and the ambient credential chain, and on provider=endpoint through a provider block carrying the base URL, with OPENAI_API_KEY in the environment. docs/codex-setup.md is the setup guide: which models each path can serve, the Responses-safe tool parsers a vLLM server needs, and the stdin redirect every codex run requires.

Codex reports token counts but no billed cost, so the harness derives cost from a new local price table, bedrock_pricing.py. A model missing from that table yields a null cost rather than a misleading zero.

Its token fields need their own convention. Codex reports input_tokens as the full prompt with the cached part as a subset, and the harness subtracts that out before recording, so the fields arrive mutually exclusive. token_accounting.py therefore declares codex's cache fields disjoint through DISJOINT_CACHE_AGENTS instead of detecting the shape from the data. Detection misfires when a codex run's cache hit rate lands near 50%: fresh input then equals cache_read + cache_write, the partition branch fires, and the total silently loses the whole cache read, halving the derived cost on the self-hosted path.

PR #162 added the agent. PR #184 made it runnable end to end and fixed the accounting.


What's changed

Documentation

  • Root README and harness-reference.md now name all five supported harnesses. The agent reference had documented only claude and pi; it now covers omp, kiro and codex, including each one's cost basis.
  • New codex-setup.md and the FAQ folder it feeds. omp and kiro-cli each had a per-harness setup page; codex's wiring lived scattered across the harness reference, the benchmark skill and a script docstring. The setup page installs codex, wires it to Bedrock and tabulates its failure modes; the FAQ's first entry, wiring codex to a model, carries the per-path recipes: the LiteLLM Responses-to-Converse bridge an open-weight Bedrock model needs, the Responses-safe tool parsers a vLLM server needs, and the direct route for openai.* models.
  • New b300-cuda-fixes.md: the one-time driver and CUDA fixes a B300 node needs before vLLM will serve (#162).
  • agent-cli-bedrock-setup.md gained the codex-on-Bedrock section: the native amazon-bedrock provider needs codex >= 0.144 and no proxy or bearer token (#162).
  • New serving guide for Kimi-K3, a 2.8T-parameter MoE with 104B active and native MXFP4 weights (#162).

Results and frontier

  • The Claude Code and combined /swe3 frontiers now include kimi-k3 at 73.92 for $8.81 a task over the 5-task v1 dataset. The run data landed before 0.1.0; this release regenerates the frontier files that had not picked it up (#182).
  • Charts, radars, decks and the vended models.json plus model-aliases.json regenerated for MiniCPM5-2B (#182).

Bug fixes

  • The codex agent was unreachable, endpoint runs silently hit Bedrock, and its token totals could drop the cache read. All three are fixed in #184, which closes #183.
  • vend/swe-router/models.json carried a source_commit that no longer existed on main, because the file was generated on a branch that GitHub then squash-merged. Restamped, which also unblocked the Test job (#184).
  • A priced Bedrock model could still record a null cost, which plots as free. _rates stripped only inference-profile prefixes, so neither the Bedrock id (qwen.qwen3-coder-30b-a3b-v1:0) nor the LiteLLM alias the harness records (...-instruct) resolved to a row. _normalize now strips a version suffix and a trailing -instruct as well (#186).
  • cost_usd indexed rates["cache_read"] unconditionally, so a model with no published cache rate raised KeyError. It now returns None when a caller reports cached tokens against such a row, rather than valuing that traffic at zero (#186).

Closed issues

Issue Title Closed by
#183 codex harness: agent is unreachable, endpoint runs hit Bedrock, and token totals can drop the cache read PR #184

Pull requests included

PR Title
#186 fix(pricing): price Qwen on Bedrock, and stop a cacheless model costing zero
#184 harness: make agent=codex runnable on the endpoint path, and fix its token accounting
#182 benchmark: MiniCPM5-2B on a single L40S -- serving guide, /swe3 results and throughput sweep
#181 benchmark: Qwen3.6-35B-A3B FP8 on a single L40S -- serving guide and results (off-frontier)
#162 feat(harness): add codex agent; Kimi-K3 guide; B300 CUDA fixes

Contributors

Thank you to everyone who contributed to this release:


Full Changelog: 0.1.0...0.2.0