Repository navigation
v3.44.2
simplicio-loop 3.44.2
bench/llm_ab/run.py --tasks 4 --hardadds a hard Python set with hidden acceptance tests outside the arm repo: coupon logic with a half-up rounding trap, a two-bug fix, a two-file refactor, and a duration parser.--turbo-reasoningkeeps the model's reasoning on for the turbo calls, so on and off can be compared. Turbo also sends a task'scontextfiles to the model. The checker accepts an absolute path.- Hard-set result: with reasoning off, turbo passed 12/12 hidden-test tasks at 5.4 s and $0.0027 per run (mean of 3). The OpenCode arm passed 11/12 at 112.6 s and $0.0122. With reasoning on, turbo passed 7/8, and one call ran away to 131k reasoning tokens. Turbo keeps reasoning off.
- Every benchmark run of 3.44.0 to 3.44.2 is archived under
bench/llm_ab/results/runs/with a summaryREADME.md, outside the release-to-release history.
PR #1374, issue #1373. Every benchmark run is archived in bench/llm_ab/results/runs/ (summary in its README.md).
Hard set, turbo reasoning off vs on (2026-09-29-hard/)
Four Python tasks with hidden tests (--tasks 4 --hard, loop at 5d11aca9).
| task | normal, 3 runs | simplicio reasoning off, 3 runs | simplicio reasoning on, 2 valid runs |
|---|---|---|---|
| 1 pricing (half-up rounding) | β β β | β β β | β β |
| 2 inventory bug fix | β β β | β β β | β β |
| 3 two-file refactor | β β β | β β β | β β |
| 4 duration parser | β β β | β β β | β β |
| passed | 11/12 | 12/12 | 7/8 |
| wall, mean | 112.6 s | 5.4 s | 18.1 s / 302.5 s |
| cost, mean | $0.0122 | $0.0027 | $0.0090 / $0.1621 |
| reasoning tokens, mean | 2,485 | 0 | 5,306 / 132,832 |
normaloff-2, task 1: half-up rounding was wrong (expected 1703 and 1712, got 1704 and 1713).- Reasoning-on
on-2, task 4: one call ran to 131,072 reasoning tokens in 294 s and returned no plan, soduration.pywas never written. on-3-http402/andon-3-rerun-http402/are invalid runs. The key ran out of credits (HTTP 402), so every call failed. They are kept only as a record.
Decision: turbo keeps reasoning off.