-
Notifications
You must be signed in to change notification settings - Fork 0
5 Performance
How fast she replies is almost entirely one question: does the model fit on your graphics card? If it does, she should be fast. If it doesn't, your GPU is too weak. consider upgrading or switching to a lower model.
So the order here is the order you should work in. Don't skip to the tuning knobs.
Every change on this page is a line in JunOS/.env followed by a restart (./JunOS/start.sh restart, or stop then start on Windows). See Manage > Changing a setting.
Send her a message so the model is actually loaded, then:
docker exec omega-ollama ollama ps # Linux
ollama ps # Windows, in a new terminalThe PROCESSOR column is the whole story.
| It says | What it means |
|---|---|
100% GPU |
she's where she should be. Go to the tuning knobs |
71% GPU / 29% CPU |
part of her is running on the processor. This is your problem. Fix that first |
100% CPU |
no card in play at all. Either you have none (fine, see CPU only) or Docker can't see it (install page) |
Two things that look like slowness but aren't:
- The first reply after a start is always slow. The weights are being read off disk into the card. Judge her on the second one.
- A long conversation gets slower to start replying. She re-reads the context before she says anything. That's prompt evaluation, and it grows with how much conversation she's holding, KV cache helps but it's not free. See context.
A model needs its weights in VRAM plus room for the context. She reserves 2 GB of working room on top of the weights and sizes the rest from what's left.
If ollama ps says anything under 100% GPU, in order:
- Close whatever else is using the card. A game, a browser with a lot of tabs, another local model. VRAM is not shared politely.
-
Drop a quant level.
Q8_0→Q6_K→Q4_K_Mon the same model is the cheapest quality you'll ever trade. See swapping the model and the table. -
Drop a size.
12B→E4B→E2B. Bigger steps down, bigger difference in what she can hold together over a long scene.
She has one self-defence here worth knowing about, so it doesn't look like a bug: Ollama decides how many layers go on the card once, when the model loads, based on whatever VRAM was free at that moment. If something else owned the card then, that bad split sticks for as long as she stays loaded. So when she notices she's partially offloaded and the weights should have fit, she evicts herself once so the next load gets an idle card. At most once every 10 minutes - if it comes back just as badly, something else genuinely has your VRAM and reloading every turn would be worse than being slow.
Reporting the card as smaller than it is. If another program permanently owns part of your VRAM, tell her so she stops planning around memory she can't have:
OMEGA_GPU_VRAM_MB=10240
How much conversation she keeps in her head. It costs VRAM - roughly 0.2 MiB per token - and it's the second biggest lever after the model itself.
She sizes it herself: card size, minus the weights, minus 2 GB of working room, then rounded down to one of 6144 / 8192 / 12288 / 16384 tokens. When the card is unknown or there's under 4 GB of room left she falls back to tiers off system RAM instead.
Override it by hand:
OMEGA_NUM_CTX=8192
Which direction do you want?
- Smaller if she's spilling onto the CPU, or if prompt evaluation before each reply has become the slow part. 6144 is still a long conversation.
- Bigger only if you have VRAM to spare and you want her to remember more of a long scene verbatim. Long-term memory doesn't live here - she consolidates chats into notes - so a huge context is not the fix for "she forgot something from last week."
Getting this wrong in the "bigger" direction is how a model that used to fit stops fitting. Check ollama ps after changing it.
The one genuinely large speedup that isn't "get a better card". Express turns it on and measures the depth for you; Custom asks and defaults to off, because whether it pays depends on your hardware.
A small assistant model guesses the next few tokens, Jun checks all of them in a single pass, and the ones she agrees with came almost free. She never says anything she wouldn't have said anyway - it changes speed, not output.
You can think about it this way: Jun has an assistant that drafts her messages, then Jun spends a tiny bit of time checking that what the assistant wrote makes sense. if it doesn't she falls back generating the text by herself
Measured on an RTX 3060 with the 12B, on prose:
| Draft depth | Speed |
|---|---|
| no drafter | 36.1 tok/s |
| 1 | 45.3 tok/s (+25%) |
| 2 | 42.0 tok/s |
| 3 | 38.2 tok/s |
| 4 | 35.4 tok/s - slower than not using it |
Deeper is not always better: every extra drafted token is one more she has to check, about 10 ms each. Where the curve turns over depends on your card, which is why there's a measuring tool instead of a recommended number:
./JunOS/mtp-autotune.sh # Linux
.\JunOS\mtp-autotune.ps1 # WindowsIt needs the stack up and the models pulled, tries every depth, and writes the fastest one into .env. Safe to re-run whenever.
It also stamps which card it measured on. Swap your GPU and the start scripts notice the stamp doesn't match what's in the box and re-measure before handing you the app. MTP_AUTOTUNE=off in .env stops that - worth knowing if you're on the llama.cpp provider, where the sweep restarts the server once per depth and that boot is not a quick one.
The drafter has to match. It must be the same Gemma size and the same QAT branch Jun was fine-tuned from. A mismatched one still runs and nothing warns you - it just guesses wrong far more often (2.10 accepted tokens per pass instead of 2.74). Left empty, the tuner derives the right one from your chat model: Jun-LoRA-12B-GGUF drafts off Jun-LoRA-12B-MTP-GGUF. Set OLLAMA_MTP= by hand only if the drafter lives somewhere else.
Gemma 4 only, and the drafter wants its own slice of VRAM. On a card that's already full this is might be worse than without. not your next move - fitting is.
These ship enabled, listed so you don't go looking for them:
| Line | What it does |
|---|---|
OLLAMA_FLASH_ATTENTION=1 |
faster attention, less memory for the same context |
OLLAMA_KV_CACHE_TYPE=q8_0 |
the context cache at 8-bit. Roughly halves what the context costs in VRAM for no quality you'd notice |
TITLE_MODEL=...Titlewen... |
a tiny separate model that names conversations. Pinned to the CPU and kept loaded on purpose, so it can never take VRAM or a GPU slot from her. Empty it to just truncate the first message instead |
| Line | What it does |
|---|---|
GPU_DEVICES=auto |
default. Orders cards biggest-VRAM-first so the largest becomes device 0 and she lands on it |
GPU_DEVICES=0,1 |
a specific card or cards, by index or UUID, when the guess is wrong |
GPU_DEVICES=all |
leave the driver's order alone |
TENSOR_PARALLEL=on |
split one model across every card. Slower per token, but fits more |
Tensor parallel is for "the model I want doesn't fit on one card", not for "make it faster". Two cards running one model is not twice the speed.
These share the machine with her, so they're on this page even though they're not the chat.
| Line | Default | Worth changing? |
|---|---|---|
TTS_DEVICE=cpu |
cpu | Leave it. CUDA costs about 2 GB of VRAM for very little synthesis speedup, and that's 2 GB taken from her |
SEP_DEVICE=auto |
auto | Song splitting is the opposite case - minutes on CPU versus seconds on a card, and it hands the VRAM straight back afterwards. Worth it |
KARAOKE=off |
on | Never going to use it? Off skips a multi-GB container image at build time on Linux |
VOICE=off |
on | Text only. Frees a few hundred MB of RAM and makes starts quicker |
STT_MODEL=base |
base |
base.en with STT_LANG=en is a bit faster and sharper if you only speak English. small / medium / large-v3 are better at other languages and slower |
She works without a card. Expect a few words per second rather than a stream - fine for a chat, patient for a long story.
- Express already picked
E2B Q4_K_Mfor you. Don't hand-pick something bigger. A 12B on a CPU is several minutes per reply. - Keep the voice on the CPU (it's already there) - both engines run in real time.
- Karaoke splitting takes minutes instead of seconds.
KARAOKE=offif you'll never use it. - MTP is mainly a GPU optimisation. Benchmark it or skip it (
OLLAMA_MTP=empty in.envturns it off).
Two ceilings, both in .env, both generous by default. Raise them on a slow CPU-only box with a big context, where a long reply can legitimately outrun them:
OMEGA_TURN_TIMEOUT_S=900 # wall clock for a whole turn, tool calls included
OMEGA_STREAM_IDLE_S=300 # hang up if nothing has arrived for this long
OMEGA_STREAM_IDLE_S has to cover a cold model load plus prompt evaluation on a full context, which is the real reason to raise it.
- A bigger model on a card it doesn't fit on. A 12B spilling onto the CPU is worse in every way than an E4B that fits.
- More system RAM, if the model is on the GPU. It's VRAM that's binding.
-
Raising
OLLAMA_NUM_PARALLEL. It's 1 because she's one person talking to one model. Parallel slots divide the context between them. - Bigger context to fix forgetfulness. Her long-term memory is consolidated notes, not the context window.