A web front-end that makes a tiny local model look intentional.
Qwen2.5-1.5B-Instruct at Q4_K_M, served from a 2 vCPU / 4 GB VPS with no GPU,
no API key and no upstream provider. The performance story isn't hidden — it's
the feature: live tokens/sec while the answer streams, time-to-first-token under
every reply, and a benchmark page that measures the machine it's running on and
plots it against reference hardware.
It also gives the model a tool belt — exact arithmetic, a clock, Wikipedia, weather, currency and a URL reader — because those are precisely the things a 1.5B model gets wrong, and they are fixable without a bigger model.
app/
main.py FastAPI: pages, SSE chat, tools, skills, benchmark, history
llm.py one Llama instance, one lock, honest metrics, grammar routing
router.py slash commands -> heuristics -> grammar-constrained routing
skills.py one-click workflows built on the tools
tools/ base (registry + GBNF) · offline · web · reader (SSRF guard)
bench.py the benchmark suite + reference data
personas.py five presets (system prompt + sampling profile)
db.py SQLite (WAL): conversations, messages, tool_calls, runs
templates/ Jinja2 — chat, bench, about
static/ hand-written CSS + ES modules, no build step
js/canvas.js WebGL2 aurora that reacts to the token stream
js/cards.js visual answers — weather, calculator tape, clock, …
scripts/bench_cli.py headless benchmark → JSON
deploy/ systemd unit + nginx site
mkdir -p models
curl -L -o models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf \
https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct-GGUF/resolve/main/qwen2.5-1.5b-instruct-q4_k_m.ggufAbout 1.0 GB. Any GGUF with a chat template works — point AURORA_MODEL_PATH at it
and update AURORA_MODEL_LABEL / AURORA_MODEL_QUANT so the UI stops lying.
python3 -m venv .venv && source .venv/bin/activate
pip install -U pip
# Build llama.cpp for this CPU. -DGGML_NATIVE=ON is worth real tokens/sec.
CMAKE_ARGS="-DGGML_NATIVE=ON" pip install --no-cache-dir llama-cpp-python
pip install -r requirements.txt
cp .env.example .envOn a 2 vCPU box the llama.cpp build takes 5–15 minutes and wants ~1 GB of free
RAM. If it gets OOM-killed, add swap first (fallocate -l 2G /swapfile), or
install a prebuilt wheel from the llama-cpp-python releases page.
uvicorn app.main:app --host 127.0.0.1 --port 8000 --workers 1One worker, always. A second worker loads a second copy of the weights and puts the box straight into swap.
AURORA_MOCK=true uvicorn app.main:app --port 8000Mock mode fakes a generator at ~11 tok/s — the whole site works, including the benchmark, so you can develop the front-end anywhere. On Windows:
$env:AURORA_MOCK="true"; .venv\Scripts\python -m uvicorn app.main:app --port 8000sudo cp deploy/aurora.service /etc/systemd/system/
sudo cp deploy/nginx.conf /etc/nginx/sites-available/aurora
sudo ln -s /etc/nginx/sites-available/aurora /etc/nginx/sites-enabled/
sudo systemctl daemon-reload && sudo systemctl enable --now aurora
sudo nginx -t && sudo systemctl reload nginxThe one thing that breaks this deploy: nginx must have
proxy_buffering offandgzip offon/api/chatand/api/bench/run. With buffering on, nginx holds every token until the response finishes, the answer lands in one lump, and every live metric on the site becomes decorative.deploy/nginx.confsets it; don't "tidy" it away.
Other deployment notes:
LimitMEMLOCK=infinityin the unit is required if you setAURORA_USE_MLOCK=true. Without it llama.cpp can't pin the weights and quietly falls back to pageable memory.MemoryMax=3Gis an OOM guard — the service restarts instead of taking the whole VPS down with it.- Keep ~2 GB of swap as a safety net. In normal operation the process should never touch it.
pip install psutilif you want the memory readouts on non-Linux hosts; on Linux they come from/proc/self/statusfor free.
Everything is AURORA_-prefixed, read from the environment or .env.
| Variable | Default | Notes |
|---|---|---|
AURORA_MODEL_PATH |
./models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf |
|
AURORA_N_CTX |
4096 |
~115 MB of KV cache at f16 |
AURORA_N_THREADS |
2 |
match the vCPU count; more threads on 2 cores loses throughput |
AURORA_N_BATCH |
256 |
prompt-eval batch; 512 raises peak RAM for little gain here |
AURORA_USE_MLOCK |
false |
true on the VPS, with the systemd unit above |
AURORA_CHAT_FORMAT |
auto |
auto uses the GGUF's own template; chatml forces Qwen's |
AURORA_HARDWARE_LABEL |
2 vCPU · 4 GB VPS |
shown throughout the UI |
AURORA_MOCK |
false |
run the site with no model present |
AURORA_MOCK_DELAY |
0.085 |
seconds between fake tokens; 0 runs the mock at full speed |
AURORA_TOOLS_ENABLED |
true |
master switch for the tool belt |
AURORA_ALLOW_OUTBOUND |
true |
false removes every network tool |
AURORA_TOOL_TIMEOUT |
10 |
seconds before a tool is abandoned |
AURORA_USER_AGENT |
— | must carry a contact URL or Wikipedia 403s |
Eight tools, no API keys anywhere:
| Tool | Network | |
|---|---|---|
calculator |
exact arithmetic via Fraction — 0.1+0.2 really is 0.3 |
— |
clock |
date and time; accepts Europe/Paris, Tokyo, new york or Japan |
— |
convert_units |
length, mass, time, data, speed, volume, temperature | — |
text_stats |
words, sentences, reading time, top terms | — |
wikipedia |
article summaries | yes |
weather |
current conditions + 5-day forecast (open-meteo) | yes |
currency |
ECB rates (Frankfurter) | yes |
read_url |
fetch a page and extract its readable text | yes |
Skills are one-click workflows on top of them — Summarise a link,
Fact-check, Compare two things, Weather brief, Do the maths,
Explain this code. Each is a dict entry in app/skills.py, not a code path.
It isn't, on its own — asked politely for JSON it produces trailing commas, prose preambles and invented keys. Four layers, cheapest first:
- Slash commands —
/calc 1247*89,/weather Paris,/read <url>. Free and certain. Press/in the composer for the palette. - Heuristics — a pasted link,
1247 * 89,15% of 200,100 EUR to USD,what time is it in Tokyo. Regex, free, and better than the model's judgement. These handle most real traffic with no model call at all. - Grammar-constrained routing — only for the Agent persona, and only when the first two miss. A short pass runs under a GBNF grammar generated from the tool registry, so malformed output is not unlikely, it is unrepresentable.
- Failure is visible — a tool that errors or times out renders as a failed step in the timeline and the model answers without it.
A routing pass costs ~2–5 s on 2 vCPU, which is why only the Agent persona pays for it. The per-step timings are printed in the UI rather than hidden.
read_url is the only tool that makes the server fetch a user-supplied
address, so it resolves the hostname first and refuses any private, loopback,
link-local, CGNAT, reserved or multicast address — including
169.254.169.254, the cloud metadata endpoint. Redirects are followed by hand,
maximum three hops, re-validating at each one. Bodies are capped at 2 MB and
only text/html / text/plain is read.
The calculator parses with ast and a node whitelist. It never calls eval,
and rejects attribute access, comprehensions, lambdas and __import__.
Set AURORA_ALLOW_OUTBOUND=false to drop every network tool from the registry,
the grammar and the UI in one move.
Wikipedia returns 403 for a User-Agent with no contact details. Set
AURORA_USER_AGENTto something including your repository URL or email.
One generation at a time. With two cores, two generations running side by side finish later than the same two run back to back. The engine takes a single lock and reports queue position over SSE, so waiting visitors see "1 request ahead" instead of a frozen page.
The lock is released by the generation thread, in its own finally — not by
the request handler. Otherwise a client disconnecting mid-answer hands the slot
to the next visitor while llama.cpp is still mid-token.
Tokens cross threads through a queue. Generation runs on a plain thread and
pushes into an asyncio.Queue, so the event loop stays free for /api/health,
cancellation, and other clients' queue counters.
The KV cache is reused. One Llama instance for the process lifetime, full
conversation sent each turn — llama.cpp prefills only the new tokens, so turn 5
starts far faster than turn 1. Creating a Llama per request would throw that
away and re-read a gigabyte of weights every time.
The chat template does the work the sampler can't. An instruct model fed raw
text behaves like a base model: it rambles, repeats, and never stops cleanly. The
usual response — clamp temperature to ~0 and bolt on a repeated-sentence
detector — makes it worse, because greedy decoding is what produces loops.
Routing every turn through <|im_start|> / <|im_end|> lets the model stop on
its own token and lets sampling sit at a normal 0.7 with a light
repeat_penalty.
The front-end has no dependencies. No build step, no CDN, no third-party
JavaScript — the markdown renderer is ~150 lines in static/js/md.js and
escapes everything before producing a tag; the icons are inline SVG, because
symbols like ⧉ and ⇹ are missing from plenty of font stacks and render as
tofu boxes on somebody else's machine.
The background is a shader that reacts to generation. static/js/canvas.js
is a hand-written WebGL2 fragment shader: domain-warped fbm noise whose
u_energy uniform spikes on every token and decays, so the aurora surges while
the model writes and settles when it stops. It renders at half resolution, caps
devicePixelRatio at 1.5, runs at 30 fps idle and 60 fps while generating, and
cancels its frame loop when the tab is hidden. Three fallbacks: no WebGL2 keeps
the CSS gradient layers, prefers-reduced-motion paints one static frame, and
a hidden tab paints nothing. All of it is client-side — the VPS pays nothing.
Run the suite from the /bench page, or headless:
python scripts/bench_cli.py --out bench.jsonThree prompt sizes (~32 / ~256 / ~1024 tokens) × 128 generated tokens, 3/3/2
repetitions, one discarded warm-up. It reports prefill tok/s, generation tok/s,
TTFT p50/p95 and peak RSS, and stores every run so /bench can chart drift.
Order-of-magnitude figures for this model and quantisation under llama.cpp. They
swing widely with memory bandwidth, AVX support, and — on shared cloud — whoever
else is on the host. The /bench page replaces the first row with what your
box actually did.
| Hardware | Threads | Prefill tok/s | Generation tok/s |
|---|---|---|---|
| 2 vCPU shared cloud | 2 | ~25–60 | ~6–12 |
| 4 vCPU dedicated, AVX2 | 4 | ~60–120 | ~12–20 |
| Desktop 8-core, DDR5 | 8 | ~200–400 | ~30–55 |
| Apple M2 (Metal) | — | ~500+ | ~45–70 |
Generating a token reads every weight once — about 1 GB at Q4_K_M — so
throughput follows memory bandwidth, not clock speed:
tokens/sec ≈ effective bandwidth ÷ model size. That's also the whole argument
for 4-bit: a quarter of the bytes moved, roughly four times the speed. Prefill is
a batched matmul and therefore compute-bound, which is why a 1024-token prompt
costs far less than 32× a 32-token one.
| Component | RAM |
|---|---|
Weights (Q4_K_M, mmap) |
~990 MB |
| KV cache @ 4096 ctx, f16 (28 layers × 2 KV heads × 128 dim = 28 KiB/token) | ~115 MB |
Compute buffers @ n_batch 256 |
~200 MB |
| Python + FastAPI + uvicorn | ~120 MB |
| Peak | ~1.4 GB of 4 GB |
A 7B model at the same quantisation is ~4.4 GB of weights alone — it wouldn't load, and if it did, every token would come off disk.
| Endpoint | |
|---|---|
POST /api/chat → SSE |
events: conversation, step, card, queue, start, token, trimmed, done, error |
GET /api/tools |
the registry, as the UI sees it |
GET /api/skills |
workflow definitions |
POST /api/cancel/{id} |
stop an in-flight generation |
GET /api/health |
model, status, RSS, queue depth, uptime |
GET/DELETE /api/conversations[/{id}] |
history |
POST /api/bench/run → SSE |
stage, rep, case, summary |
GET /api/bench/results |
latest summary + history + reference data |
curl -N -X POST localhost:8000/api/chat \
-H 'Content-Type: application/json' \
-d '{"message":"count to twenty"}'Tokens must appear one at a time. If they arrive all at once, something between you and uvicorn is buffering — see the nginx note above.
It's a 1.5B model at 4 bits. It gets arithmetic wrong, invents citations, loses long reasoning chains, and knows nothing recent. It's good at short explanations, quick drafts, summarising pasted text, and simple code — in a couple of seconds, without a byte leaving the machine.