Launch Claude Code with multiple backends (Anthropic, OpenRouter, Ollama, NVIDIA NIM, LM Studio, llama.cpp).
- Multiple backends - Anthropic, OpenRouter, Ollama, NVIDIA NIM, LM Studio, llama.cpp
- OAuth login - authenticate with
claude-launcher login - Model picker - searchable model selection
- Exacto support - auto-uses
:exactovariants for better tool calling - Role models - configure different models for sonnet/opus/haiku tasks
- New model alerts - notifies when new models are available
- Local backends - run fully offline via Ollama, LM Studio, or llama.cpp
npm install -g claude-launcher
# or
pnpm add -g claude-launcher
# or
yarn global add claude-launcher
# or
bun add -g claude-launcherRequires Claude Code installed.
claude-launcher # launch with saved settings
claude-launcher login # authenticate with OpenRouter
claude-launcher logout # clear stored credentials
claude-launcher -b # pick backend and model
claude-launcher -a # use Anthropic backend
claude-launcher -o # use OpenRouter backend
claude-launcher -l # use Ollama backend (local)
claude-launcher -n # use NVIDIA NIM backend
claude-launcher -s # use LM Studio backend (local)
claude-launcher --llamacpp # use llama.cpp backend (local)
claude-launcher -k # use Kimi backend (Kimi Code subscription)
claude-launcher -x # use Mixed backend (Anthropic native + any backend per role slot)
claude-launcher -- --resume # pass args to claude- Anthropic - standard Claude Code, no extra config
- OpenRouter - any model via OpenRouter; OAuth login or
OPENROUTER_API_KEY - Codex - run Claude Code on a ChatGPT subscription (reuses the
codexCLI login) - Grok - run Claude Code on an X/SuperGrok subscription (reuses the
grokCLI login) - Kimi - run Claude Code on a Kimi Code subscription (reuses the
kimiCLI login) - Mixed - any backend per role slot in one session: main/opus/sonnet/haiku each stay on the Claude Pro/Max subscription or route to Codex, Grok, Kimi, OpenRouter, Ollama, NIM, LM Studio, or llama.cpp. Slot models use a
backend:modelprefix and requests are routed per model, so/model codex:gpt-5.6-solor/model ollama:qwen3mid-session works. - Ollama - local models, auto-filtered to tool-capable ones
- NVIDIA NIM - cloud (
NVIDIA_API_KEY) or self-hosted endpoints - LM Studio - local models via the LM Studio server (host must include
/v1, e.g.http://localhost:1234/v1) - llama.cpp - a local or LAN
llama-serverinstance; setup prompts for the server URL (server root, e.g.http://localhost:8080orhttp://my-box.local:8080)
NIM and LM Studio run through an in-process Anthropic-to-OpenAI translation proxy. Ollama and llama.cpp speak the Anthropic Messages API natively, so Claude Code talks to them directly.
- Run
claude-launcher -b - Select a backend
- Provide credentials if the backend needs them
- Pick a model
- Optionally configure role models (sonnet/opus/haiku)
Settings stored at ~/.config/claude-launcher/config.json:
- Backend preference
- Selected models (main, sonnet, opus, haiku)
- API key (if logged in via OAuth)
OPENROUTER_API_KEY- fallback if not logged in via OAuthNVIDIA_API_KEY- NIM cloud API key
MIT