You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
Kimi now signs in at auth.kimi.ai and calls api.kimi.ai, where Kimi
moved. Both hosts share one sign-in, so an existing login keeps working
with no new cc-proxy kimi auth login. kimi.oauthHost and kimi.baseUrl still override them.
cc-proxy shell install makes plain claude go through cc-proxy. It adds
one line to ~/.zshrc, ~/.bashrc or ~/.bash_profile, or a function
file for fish. After that, cc-proxy off and cc-proxy on switch every
terminal at once, and claude --resume <id> keeps working. The line sets
no environment variables, so other programs never see the proxy. cc-proxy shell uninstall removes it.
cc-proxy claude starts Claude Code on the proxy. It starts the proxy
first if it isn't running, then passes every argument to claude, so --resume, --worktree, -p and a pasted resume command work as usual.
The proxy's address goes in a --settings flag for that session, so no
settings file changes. Choose the model it starts on with cc-proxy config set claude.model <id>, and the background model with claude.fastModel.
The serve banner and the monitor's setup panel now point to cc-proxy claude. Their variable lists use ANTHROPIC_DEFAULT_HAIKU_MODEL,
which replaced ANTHROPIC_SMALL_FAST_MODEL in Claude Code, and add CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1.
cc-proxy usage shows how much of your Codex, Kimi and OpenCode Go plans
you have used, and when each limit resets. Name one provider to see only
that one: cc-proxy usage kimi. Add --json for scripts.
cc-proxy opencode usage moved to cc-proxy usage opencode. The old
command is gone. With --json, OpenCode Go's reply now sits under an opencode key.
cc-proxy models no longer lists five Codex models that ChatGPT
accounts can't use: gpt-5.2, gpt-5.3-codex, gpt-5.3-codex-spark, gpt-5.4
and gpt-5.4-mini. The backend answers each with a 400. On the OpenAI
routes they now fail at once with the list of models that work.
GLM and OpenCode Go's Anthropic models (MiniMax, Qwen) now send Claude
Code only whole stream events. When an upstream read ended halfway through
an event and an error came next, the error was stuck onto the half event,
so Claude Code could read neither, and text from the same read was lost.
OpenCode Go's GPT, Grok and Muse Spark models now pass on what failed when
the upstream fails mid-answer. A rate limit used to reach Claude Code as a
generic "stream is invalid" error, without the upstream's message, and the
text that came just before it was dropped.
Cursor now tells a spent quota from a short rate limit. Both arrive as a
429, so Claude Code kept retrying a spent quota, which a few seconds of
waiting can't fix. A spent quota now comes back with x-should-retry: false,
which stops the retries, whether Cursor reports it as an HTTP error or at
the end of its answer.
A Cursor 429 now passes on Cursor's reason and its own retry-after.
Before, Claude Code got "Cursor upstream error" and a 5-second wait that
Cursor never sent.
A finished answer no longer fails when an upstream read ends halfway
through what comes after the end of the answer, such as data: [DONE] or
OpenCode Go's closing metadata. This affected Kimi, GLM and OpenCode Go.
Claude Code got an error instead of the answer it had already been sent.