Skip to content

5 Performance

effx_ edited this page Sep 17, 2026 · 1 revision

Performance

How fast she replies is almost entirely one question: does the model fit on your graphics card? If it does, she should be fast. If it doesn't, your GPU is too weak. consider upgrading or switching to a lower model.

So the order here is the order you should work in. Don't skip to the tuning knobs.

Every change on this page is a line in JunOS/.env followed by a restart (./JunOS/start.sh restart, or stop then start on Windows). See Manage > Changing a setting.


1. Measure first

Send her a message so the model is actually loaded, then:

docker exec omega-ollama ollama ps    # Linux
ollama ps                             # Windows, in a new terminal

The PROCESSOR column is the whole story.

It says What it means
100% GPU she's where she should be. Go to the tuning knobs
71% GPU / 29% CPU part of her is running on the processor. This is your problem. Fix that first
100% CPU no card in play at all. Either you have none (fine, see CPU only) or Docker can't see it (install page)

Two things that look like slowness but aren't:

  • The first reply after a start is always slow. The weights are being read off disk into the card. Judge her on the second one.
  • A long conversation gets slower to start replying. She re-reads the context before she says anything. That's prompt evaluation, and it grows with how much conversation she's holding, KV cache helps but it's not free. See context.

2. Make her fit

A model needs its weights in VRAM plus room for the context. She reserves 2 GB of working room on top of the weights and sizes the rest from what's left.

If ollama ps says anything under 100% GPU, in order:

  1. Close whatever else is using the card. A game, a browser with a lot of tabs, another local model. VRAM is not shared politely.
  2. Drop a quant level. Q8_0 → Q6_K → Q4_K_M on the same model is the cheapest quality you'll ever trade. See swapping the model and the table.
  3. Drop a size. 12B → E4B → E2B. Bigger steps down, bigger difference in what she can hold together over a long scene.

She has one self-defence here worth knowing about, so it doesn't look like a bug: Ollama decides how many layers go on the card once, when the model loads, based on whatever VRAM was free at that moment. If something else owned the card then, that bad split sticks for as long as she stays loaded. So when she notices she's partially offloaded and the weights should have fit, she evicts herself once so the next load gets an idle card. At most once every 10 minutes - if it comes back just as badly, something else genuinely has your VRAM and reloading every turn would be worse than being slow.

Reporting the card as smaller than it is. If another program permanently owns part of your VRAM, tell her so she stops planning around memory she can't have:

OMEGA_GPU_VRAM_MB=10240

3. Context window

How much conversation she keeps in her head. It costs VRAM - roughly 0.2 MiB per token - and it's the second biggest lever after the model itself.

She sizes it herself: card size, minus the weights, minus 2 GB of working room, then rounded down to one of 6144 / 8192 / 12288 / 16384 tokens. When the card is unknown or there's under 4 GB of room left she falls back to tiers off system RAM instead.

Override it by hand:

OMEGA_NUM_CTX=8192

Which direction do you want?

  • Smaller if she's spilling onto the CPU, or if prompt evaluation before each reply has become the slow part. 6144 is still a long conversation.
  • Bigger only if you have VRAM to spare and you want her to remember more of a long scene verbatim. Long-term memory doesn't live here - she consolidates chats into notes - so a huge context is not the fix for "she forgot something from last week."

Getting this wrong in the "bigger" direction is how a model that used to fit stops fitting. Check ollama ps after changing it.


4. Speculative decoding (MTP)

The one genuinely large speedup that isn't "get a better card". Express turns it on and measures the depth for you; Custom asks and defaults to off, because whether it pays depends on your hardware.

A small assistant model guesses the next few tokens, Jun checks all of them in a single pass, and the ones she agrees with came almost free. She never says anything she wouldn't have said anyway - it changes speed, not output.

You can think about it this way: Jun has an assistant that drafts her messages, then Jun spends a tiny bit of time checking that what the assistant wrote makes sense. if it doesn't she falls back generating the text by herself

Measured on an RTX 3060 with the 12B, on prose:

Draft depth Speed
no drafter 36.1 tok/s
1 45.3 tok/s (+25%)
2 42.0 tok/s
3 38.2 tok/s
4 35.4 tok/s - slower than not using it

Deeper is not always better: every extra drafted token is one more she has to check, about 10 ms each. Where the curve turns over depends on your card, which is why there's a measuring tool instead of a recommended number:

./JunOS/mtp-autotune.sh      # Linux
.\JunOS\mtp-autotune.ps1     # Windows

It needs the stack up and the models pulled, tries every depth, and writes the fastest one into .env. Safe to re-run whenever.

It also stamps which card it measured on. Swap your GPU and the start scripts notice the stamp doesn't match what's in the box and re-measure before handing you the app. MTP_AUTOTUNE=off in .env stops that - worth knowing if you're on the llama.cpp provider, where the sweep restarts the server once per depth and that boot is not a quick one.

The drafter has to match. It must be the same Gemma size and the same QAT branch Jun was fine-tuned from. A mismatched one still runs and nothing warns you - it just guesses wrong far more often (2.10 accepted tokens per pass instead of 2.74). Left empty, the tuner derives the right one from your chat model: Jun-LoRA-12B-GGUF drafts off Jun-LoRA-12B-MTP-GGUF. Set OLLAMA_MTP= by hand only if the drafter lives somewhere else.

Gemma 4 only, and the drafter wants its own slice of VRAM. On a card that's already full this is might be worse than without. not your next move - fitting is.


5. Already on for you

These ship enabled, listed so you don't go looking for them:

Line What it does
OLLAMA_FLASH_ATTENTION=1 faster attention, less memory for the same context
OLLAMA_KV_CACHE_TYPE=q8_0 the context cache at 8-bit. Roughly halves what the context costs in VRAM for no quality you'd notice
TITLE_MODEL=...Titlewen... a tiny separate model that names conversations. Pinned to the CPU and kept loaded on purpose, so it can never take VRAM or a GPU slot from her. Empty it to just truncate the first message instead

6. Two graphics cards

Line What it does
GPU_DEVICES=auto default. Orders cards biggest-VRAM-first so the largest becomes device 0 and she lands on it
GPU_DEVICES=0,1 a specific card or cards, by index or UUID, when the guess is wrong
GPU_DEVICES=all leave the driver's order alone
TENSOR_PARALLEL=on split one model across every card. Slower per token, but fits more

Tensor parallel is for "the model I want doesn't fit on one card", not for "make it faster". Two cards running one model is not twice the speed.


7. Voice, songs and the microphone

These share the machine with her, so they're on this page even though they're not the chat.

Line Default Worth changing?
TTS_DEVICE=cpu cpu Leave it. CUDA costs about 2 GB of VRAM for very little synthesis speedup, and that's 2 GB taken from her
SEP_DEVICE=auto auto Song splitting is the opposite case - minutes on CPU versus seconds on a card, and it hands the VRAM straight back afterwards. Worth it
KARAOKE=off on Never going to use it? Off skips a multi-GB container image at build time on Linux
VOICE=off on Text only. Frees a few hundred MB of RAM and makes starts quicker
STT_MODEL=base base base.en with STT_LANG=en is a bit faster and sharper if you only speak English. small / medium / large-v3 are better at other languages and slower

CPU only

She works without a card. Expect a few words per second rather than a stream - fine for a chat, patient for a long story.

  • Express already picked E2B Q4_K_M for you. Don't hand-pick something bigger. A 12B on a CPU is several minutes per reply.
  • Keep the voice on the CPU (it's already there) - both engines run in real time.
  • Karaoke splitting takes minutes instead of seconds. KARAOKE=off if you'll never use it.
  • MTP is mainly a GPU optimisation. Benchmark it or skip it (OLLAMA_MTP= empty in .env turns it off).

When a turn takes too long

Two ceilings, both in .env, both generous by default. Raise them on a slow CPU-only box with a big context, where a long reply can legitimately outrun them:

OMEGA_TURN_TIMEOUT_S=900    # wall clock for a whole turn, tool calls included
OMEGA_STREAM_IDLE_S=300     # hang up if nothing has arrived for this long

OMEGA_STREAM_IDLE_S has to cover a cold model load plus prompt evaluation on a full context, which is the real reason to raise it.


Things that won't help

  • A bigger model on a card it doesn't fit on. A 12B spilling onto the CPU is worse in every way than an E4B that fits.
  • More system RAM, if the model is on the GPU. It's VRAM that's binding.
  • Raising OLLAMA_NUM_PARALLEL. It's 1 because she's one person talking to one model. Parallel slots divide the context between them.
  • Bigger context to fix forgetfulness. Her long-term memory is consolidated notes, not the context window.

Clone this wiki locally