Skip to content

v0.1.49

Choose a tag to compare

@gusnips gusnips released this 02 Oct 01:07
· 9 commits to main since this release
eda86c4
  • Kimi now signs in at auth.kimi.ai and calls api.kimi.ai, where Kimi
    moved. Both hosts share one sign-in, so an existing login keeps working
    with no new cc-proxy kimi auth login. kimi.oauthHost and
    kimi.baseUrl still override them.
  • cc-proxy shell install makes plain claude go through cc-proxy. It adds
    one line to ~/.zshrc, ~/.bashrc or ~/.bash_profile, or a function
    file for fish. After that, cc-proxy off and cc-proxy on switch every
    terminal at once, and claude --resume <id> keeps working. The line sets
    no environment variables, so other programs never see the proxy.
    cc-proxy shell uninstall removes it.
  • cc-proxy claude starts Claude Code on the proxy. It starts the proxy
    first if it isn't running, then passes every argument to claude, so
    --resume, --worktree, -p and a pasted resume command work as usual.
    The proxy's address goes in a --settings flag for that session, so no
    settings file changes. Choose the model it starts on with
    cc-proxy config set claude.model <id>, and the background model with
    claude.fastModel.
  • The serve banner and the monitor's setup panel now point to
    cc-proxy claude. Their variable lists use ANTHROPIC_DEFAULT_HAIKU_MODEL,
    which replaced ANTHROPIC_SMALL_FAST_MODEL in Claude Code, and add
    CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1.
  • cc-proxy usage shows how much of your Codex, Kimi and OpenCode Go plans
    you have used, and when each limit resets. Name one provider to see only
    that one: cc-proxy usage kimi. Add --json for scripts.
  • cc-proxy opencode usage moved to cc-proxy usage opencode. The old
    command is gone. With --json, OpenCode Go's reply now sits under an
    opencode key.
  • cc-proxy models no longer lists five Codex models that ChatGPT
    accounts can't use: gpt-5.2, gpt-5.3-codex, gpt-5.3-codex-spark, gpt-5.4
    and gpt-5.4-mini. The backend answers each with a 400. On the OpenAI
    routes they now fail at once with the list of models that work.
  • GLM and OpenCode Go's Anthropic models (MiniMax, Qwen) now send Claude
    Code only whole stream events. When an upstream read ended halfway through
    an event and an error came next, the error was stuck onto the half event,
    so Claude Code could read neither, and text from the same read was lost.
  • OpenCode Go's GPT, Grok and Muse Spark models now pass on what failed when
    the upstream fails mid-answer. A rate limit used to reach Claude Code as a
    generic "stream is invalid" error, without the upstream's message, and the
    text that came just before it was dropped.
  • Cursor now tells a spent quota from a short rate limit. Both arrive as a
    429, so Claude Code kept retrying a spent quota, which a few seconds of
    waiting can't fix. A spent quota now comes back with x-should-retry: false,
    which stops the retries, whether Cursor reports it as an HTTP error or at
    the end of its answer.
  • A Cursor 429 now passes on Cursor's reason and its own retry-after.
    Before, Claude Code got "Cursor upstream error" and a 5-second wait that
    Cursor never sent.
  • A finished answer no longer fails when an upstream read ends halfway
    through what comes after the end of the answer, such as data: [DONE] or
    OpenCode Go's closing metadata. This affected Kimi, GLM and OpenCode Go.
    Claude Code got an error instead of the answer it had already been sent.