Skip to content

v3.44.1

Choose a tag to compare

@wesleysimplicio wesleysimplicio released this 29 Sep 10:12
· 206 commits to main since this release
c021887

simplicio-loop 3.44.1

The turbo path got faster and cheaper, and two 3.44.0 packaging defects are fixed.

  • Turbo calls pin the arm's OpenRouter session (x-session-id) and switch reasoning off ("reasoning": {"enabled": false}).
  • The wave no longer sleeps 3 s after the first call.
  • Every call records latency_s and provider.
  • bench/llm_ab/run.py --tasks 10 --independent measures the fan-out without the standard dependency chain.
  • The bench docs now state that in turbo mode the simplicio arm is the loop engine calling OpenRouter directly, not OpenCode.
  • simplicio-loop update repairs an install that still carries the standalone simplicio-cli / simplicio-mapper.
  • simplicio-py doctor --upgrade no longer reinstalls simplicio-mapper from PyPI.

PR #1372, issue #1371.

Benchmark after the change (turbo, deepseek-v4.1-flash, two runs A/B + one independent run)

Every task passed in both arms and every command exited 0. Cost is the summed per-task cost (settled ledger when available, otherwise computed from tokens).

tasks run normal (OpenCode) simplicio (turbo engine)
1 A 17.6 s · $0.00230 · cache 60% 2.3 s · $0.00142 · cache 0%
1 B 17.1 s · $0.00068 · cache 96% 2.1 s · $0.00142 · cache 0%
4 A 65.0 s · $0.01199 · cache 75% 9.3 s · $0.00237 · cache 75%
4 B 87.3 s · $0.00399 · cache 84% 6.1 s · $0.00258 · cache 75%
10 A 161.9 s · $0.02636 · cache 71% 11.4 s · $0.00455 · cache 78%
10 B 165.5 s · $0.01551 · cache 83% 12.7 s · $0.00306 · cache 89%
10 independent C 201.6 s · $0.01947 · cache 86% 5.3 s · $0.00308 · cache 87%

Simplicio arm against 3.44.0 (mean of two runs each):

tasks wall cost reasoning tokens
1 15.5 s → 2.2 s (−86%) $0.00280 → $0.00142 (−49%) 1,103 → 0
4 17.0 s → 7.7 s (−55%) $0.00507 → $0.00247 (−51%) 820 → 0
10 54.6 s → 12.0 s (−78%) $0.00467 → $0.00381 (−19%) 166 → 0
  • Per-call latency: 0.8–1.4 s, all on one provider (Together). Reasoning was 0 on every call.
  • Independent set: the same ten pages without the chain ran in 5.3 s, against 11.4–12.7 s chained.
  • One task: the loop is ~8× faster. On cost it wins only while the OpenCode prompt is not fully cached ($0.00142 vs $0.00230 at 60% cache; $0.00068 at 96% cache).

Results: bench/llm_ab/results/2026-09-29-790061e2-t{1,4,10,10-ind}.json. The B-run JSON files are attached to the release.