v3.7.0
TL;DR — Claude Sonnet 5.5 arrives (Opus 5.5 stays the default), Claude replies that start with thinking no longer come back empty, and a request Claude declines now says so instead of going quiet.
Added
- Claude Sonnet 5.5 (
claude-sonnet-5-5), Anthropic's new Sonnet: $2/$10
like Sonnet 5, 1M context, 128K output, and thinking on by default. Codeep
leaves it at least 32K reply tokens (64K at Max), as it does for Opus 5.5,
because its thinking shares that limit. Through OpenRouter it is
anthropic/claude-sonnet-5.5. Sonnet 5 stays in the picker, and a setting
pinned to it keeps it. Opus 5.5 stays the Anthropic default.
Fixed
- Claude replies that began with thinking could come back empty wherever
Codeep did not stream them: conversation summaries (compaction,/recall),
session titles, the review hook, the text-tool fallback and the task planner,
which then dropped its plan. Codeep read only the first block of the reply,
and on Opus 5.5 and Sonnet 5/5.5 that is often the thinking. It now reads
every text block. - A request Claude declines now says so: "Claude declined this request
(category: cyber)." It used to be an empty reply, and in agent mode the empty
turn was sent again. A declined agent turn runs none of its tool calls.
Where Codeep uses a reply itself rather than showing it, a decline counts as
no reply: no session title, conversation summary or/recallrecap is made
from it,/planfails instead of keeping it as the plan/gowould run, a
skill stops at that step instead of passing it on (/stashwould have
stashed under it), and an MCP server's sampling request gets an error
instead of an empty answer. /compactno longer throws the conversation away when the summary fails.
An empty summary replaced the earlier messages anyway, leaving the compaction
marker over nothing. The history now stays as it was, and Codeep says why
nothing was compacted.- The context meter sizes OpenRouter models by their own window.
anthropic/claude-sonnet-5.5,anthropic/claude-opus-5.5,z-ai/glm-5.3
and the other OpenRouter ids of models Codeep knows were all sized at the
128K default, so a long run warned "Context at 80% of 128k window" at about
100K tokens of a 1M window. They now get the model's window, or OpenRouter's
where it is smaller (GPT-5.5: 1.05M).