Skip to content

Local AI Setup

Luke Benko edited this page Oct 8, 2026 · 1 revision

Local AI Setup (LM Studio and Ollama)

Run your co-pilot's brain on your own graphics card. The short version: install the app, download a model, start its server, and pick it in AICS -> Settings -> AI backend.

Why go local

  • Free. No subscription, no per-message cost.
  • Offline. Once the model is downloaded, no internet needed.
  • Private. Your chats never leave your PC.
  • No API key. Nothing to sign up for, nothing to paste.

The catch: KSP and the model share one graphics card. See Which model? below.

LM Studio

  1. Install LM Studio and start it.
  2. In the search (Discover) tab, search for qwen3.5-9b (or a model from the table below) and download it.
  3. Load it with the context length set to 16384 in the load settings. The bridge's instructions and tool list don't fit in a small context.
  4. Open the Developer tab and start the server. It runs on http://localhost:1234.
  5. In KSP: AICS -> Settings -> AI backend -> LM Studio, or type /ai local in the chat. It takes effect on your next message, no restart.

The bridge uses whichever model is loaded. If nothing is loaded, it starts the server and loads qwen/qwen3.5-9b with a 16k context on its own (through LM Studio's lms tool). /model lists your models, and /model gemma switches (LM Studio swaps on the next message, about 30 s).

Ollama

  1. Install Ollama. On Windows it keeps running in the background after install; if it isn't, start the Ollama app or run ollama serve.

  2. Download a model in a command prompt: ollama pull qwen3.5:9b

  3. Raise the context: in the Ollama app's settings, set the context length slider to 16k. On cards under 24 GB Ollama defaults to 4k, which is too small for the bridge.

  4. In KSP: AICS -> Settings -> AI backend -> Ollama, or type /ai ollama.

  5. Set the model name. Without one, the bridge uses the first model Ollama lists. Type /model qwen3.5:9b in the chat for this session, or make it stick by adding this line to the .env file in the bridge folder and restarting the bridge (restart KSP, or run powershell -File tools\restart_bridge.ps1):

    OLLAMA_MODEL=qwen3.5:9b
    

[COMMS] Bill (Eng): A co-pilot that lives in the computer. Finally, someone who won't eat the snacks.

Which model?

From our own tests: an RTX 3080 with 10 GB and KSP running alongside, eight test orders per model (full results). Treat the card sizes as a rough guide; KSP itself takes a good chunk.

Graphics card Model LM Studio Ollama
10 GB or more Qwen 3.5 9B (Q4). The default: right tool 8/8, the most careful with burns qwen/qwen3.5-9b qwen3.5:9b
about 8 GB Gemma 4 E4B. Also 8/8, a bit faster, about 1.3 GB lighter in LM Studio google/gemma-4-e4b gemma4:e4b
tighter, or you want speed Qwen3 4B Instruct 2507. Fast, 6/8: fine for status, science and naming; don't trust it with numbers or landings qwen3-4b-instruct-2507 qwen3:4b-instruct

On Ollama, gemma4:e4b downloads at about the same size as qwen3.5:9b, so there the 4B is the lighter pick. Skip coder models and tiny vision models: they can't make proper tool calls, so they chat nicely and fly nothing.

[INTERCOM] Bob (Sci): It lives on our own computer? Then it's crew. I'm making it a name tag.

Quick troubleshooting

  • "Can't reach Ollama" or "Is LM Studio's server running?" The server isn't running. LM Studio: Developer tab, start the server. Ollama: start the app or run ollama serve.
  • Wrong port. AICS expects http://localhost:1234/v1 (LM Studio) and http://localhost:11434/v1 (Ollama). If you changed it, add LMSTUDIO_URL=... or OLLAMA_URL=... to .env and restart the bridge.
  • Slow replies or KSP stutters. The model is too big next to KSP. Pick a smaller one from the table, or try a cloud backend (see Installation).
  • It chats but never does anything. The context is too small (step 3 above) or the model can't call tools.

Back to Installation.

Clone this wiki locally