v0.0.302.15
Codegraff v0.0.302.15
DeepSeek now thinks at its low level at graff's default effort, so a
default turn gets to its first tool call sooner. Measured on DeepSeek V4 Pro
against the previous release:
Evals
DeepSeek V4 Pro at the default effort
v0.0.302.14 and this release on DeepSeek's own API with one key, interleaved
over 11 coding and 10 MCP tasks, 2 runs each:
| Pass | Wall per task | Coding tasks | MCP tasks | Output tokens per task | |
|---|---|---|---|---|---|
| v0.0.302.14 | 42/42 | 21.6s | 6.8s | 37.8s | 1,941 |
| v0.0.302.15 | 41/42 | 15.9s | 6.2s | 26.5s | 1,235 |
This release took 26% less time per task and wrote 36% fewer output tokens,
most of it on the MCP tasks. It missed one run: on regex-count it counted
lines with grep -c '^ERROR', which also matches the ERRORS_TOTAL line the
prompt rules out. An earlier round, with a build that sends the same request,
measured 11.9s against 24.2s per task.
On the sub-agent suites both passed 20 of 20 runs, in 37.7s per task against
61.9s: 52.9s against 85.7s when the tasks ask for sub-agents and 22.6s against
38.0s when they do not.
Turning thinking off altogether was faster again but missed more: one coding
run and 5 of 20 sub-agent runs, against none and 2 of 20 for low in the
same rounds. /effort low still sends thinking off for anyone who wants that
trade, and /effort high thinks longer.
Full results.
DeepSeek thinks at its low level by default
graff's default effort is medium, and graff sent it to DeepSeek as
reasoning_effort: medium with thinking on. DeepSeek documents low, high
and max, and on V4 Pro the first request of a turn planned the whole job
before its first tool call. The default now sends low with thinking still
enabled. /effort low still turns thinking off, high and above are
unchanged, and the flash models' default stays thinking off. The picker,
status line and ACP still call the default Medium
(ADR 0237, #1456).