▄███████████▄
▄███████████████▄
███ ▄ █████ ▄ ███ ▄ J A N A S
███▄▄▄█████▄▄▄████
▀█████▄▀▀▀▄█████▀ ▀ a family of tools for artificial intelligence, in C
█████████████
█ █ █ █ █ █ █ █
▀ ▀ ▀ ▀ ▀ ▀ ▀ ▀
Janas is a family of tools for artificial intelligence, written in C with no dependencies beyond the C library. Its first component, Janas-LLM, runs large mixture-of-experts language models on ordinary computers, with or without a GPU: the weights stay on disk, the experts each token needs are streamed from an NVMe SSD, and the ones that keep coming back stay in memory. An 80-billion parameter model runs, and answers, on a laptop with 32 GB of RAM.
This is an experiment, not a product.
| Stage | Early development. Formats, file layouts, the public API and the command line change without notice, and have changed several times a week. |
| Stability | Not verified. The engine is deliberately hard on the machine: it fills the free memory with its expert cache and reads the disk at full speed with O_DIRECT. On a machine with less headroom than it reckons, the desktop can be pushed into swap and become slow or unresponsive. Using a GPU has, once, caused a driver reset that took the desktop with it. |
| Data | It writes only in ~/.cache/janas (what it learns about your machine) and where you tell it to put a converted model. It does not touch the model files it reads. |
| Hardware | Developed and run on one machine (below). On anything else it is untested: it may be slower, it may refuse the model, it may misbehave. |
If something goes wrong, Ctrl-C stops a reply and Ctrl-D leaves the chat;
--cache <GiB> puts a hard limit on the memory it takes, and JANAS_GPU=0
keeps it off the GPU entirely.
Everything in this repository was written and measured on:
- Linux Debian 13, x86-64
- Intel Core Ultra 9 185H (6 performance cores, 8 efficiency, 2 low-power; AVX2, FMA, F16C, AVX-VNNI; integrated Arc GPU)
- 32 GB of RAM
- 1 TB NVMe SSD (about 5.7 GB/s reading)
Everything the engine does is chosen from what it measures, so another machine
will make other choices — that part is meant to travel. What has never been
tried is the rest: other CPUs (especially without AVX-VNNI, or AMD), discrete
GPUs, slower disks, more or less memory. If you run it somewhere else, the
report from janas-bench is the single most useful thing you can send back
(see Helping it run on more hardware).
- Models: Qwen3.5 and Qwen3.6 (
qwen35moe, e.g. Qwen3.6-35B-A3B), Qwen3-Next (qwen3next, e.g. Qwen3-Next-80B-A3B-Instruct and Qwen3-Coder-Next), Qwen3 MoE (qwen3moe, e.g. Qwen3-30B-A3B) and the dense Qwen3 models (qwen3), from the usual Q4_K_M GGUF files. A dense feed-forward is a mixture with one expert and no router, so the same path serves both and there is no second engine to keep in step. Of the dense ones Qwen3-4B and Qwen3-0.6B are the two that have been run here. Nearly everything here was measured on Qwen3-Next-80B-A3B-Instruct: it is the one the engine was shaped around, and the one whose numbers are quoted below. The others have been converted, checked against a reference and used, but far less. - Reasoning models: the reasoning is shown apart, in grey, and can be turned off.
janas-chat: a terminal chat with a line editor of its own, history, colours, a status line and a context that slides instead of ending.libjanas_llm: a C library (include/janas/llm.h) and a FreeBASIC binding (include/janas/llm.bi).janas-server: the model behind an HTTP API that follows OpenAI's, so that clients written for it can use a model running on this machine: chat and completions with tools, JSON output held to a schema and log-probabilities, the Responses API with its conversations, embeddings and moderations, and a chat page of its own to try it from a browser (below).janas-bench: measures the machine, fills its profile and writes a report.- Faster replies from drafts the model verifies, so the text is exactly the
one it would have written: from the model's own multi-token prediction block
where it has one, from the conversation where it does not, or from a small
model of the same family given with
--draft. - Optional GPU (Vulkan 1.3, integrated or discrete), used only where the engine measures that it helps. Results are identical with and without it.
On the machine above, with nothing else running, janas-bench reports this:
| Model | reading a prompt | writing a reply | with drafts |
|---|---|---|---|
| Qwen3-4B (dense, 2.5 GB) | 81 tok/s | 25 tok/s | see --draft below |
| Qwen3-30B-A3B (18.6 GB) | 85 tok/s | 30 tok/s | — |
| Qwen3.6-35B-A3B (22.3 GB) | 70 tok/s | 23 tok/s | 29 tok/s |
| Qwen3-Next-80B-A3B (48.4 GB) | 60 tok/s | 23 tok/s | 30 tok/s |
Every figure is the mean of three rounds, not a best round, for whichever thread and GPU setting came out fastest of the six it tries. The rounds run in alternating order so that the machine warming up and the cache filling weigh on each setting alike. Expect a few per cent either way between runs, and rather less than these on a first run, while the cache is still filling.
The last column is the model's own prediction block guessing the next tokens
and the model checking them, which is the same text at a higher rate — the
whole of it is explained under --mtp. Qwen3-30B-A3B has
no such block, and no dense model has one.
Try it on an idle machine first. The expert cache takes the memory that is free when the model is opened, and what is left of the model is read from disk while it answers, so everything else running takes its share. It degrades gently rather than breaking: measured with a browser and an editor open, the same model kept about three quarters of its speed and waited five times longer on the disk. But the first thing you see should be the machine's real speed, and any measurement you mean to send to others has to be taken with the machine to itself.
- INSTALL.md — what the build needs, how to prepare a model, and the first run
- ChangeLog.md — what changed, newest first
- MODELS.md — what a converted model's licence is, and what may be redistributed
- CONTRIBUTING.md — how to help, and the sign-off
- LICENSE — GNU GPL, version 3 or later
sudo apt install build-essential # Debian/Ubuntu
sudo apt install libvulkan-dev glslang-tools vulkan-tools # optional, GPU
./build.sh # everything into bin/x86_64-linux/
./build.sh release test # and run the testsJanas reads its own format, .jns, converted from a Q4_K_M GGUF. Keep both
in models/ next to the sources — the tools take a path, so anywhere works, but
that is where these examples put them, and where the project keeps its own.
| Model | GGUF converted here | GGUF | .jns |
Memory it likes |
|---|---|---|---|---|
| Qwen3-Next-80B-A3B-Instruct — most of this engine was measured on it | Qwen | 48.4 GB | 48.4 GB | 32 GB |
| Qwen3.6-35B-A3B — the easiest to start with | bartowski | 22.3 GB | 22.3 GB | 16-32 GB |
| Qwen3-Coder-Next | Qwen | 48.4 GB | 48.4 GB | 32 GB |
| Qwen3-30B-A3B | Qwen | 18.6 GB | 18.6 GB | 16 GB |
| Qwen3-4B — dense, and the smallest here that answers well | Qwen | 2.5 GB | 2.5 GB | 8 GB |
Qwen3-0.6B — not to talk to: to draft for a dense one (--draft) |
unsloth | 0.40 GB | 0.40 GB | with the model above |
Qwen3-Embedding-0.6B — not to talk to: for embeddings (janas-server --embedding-model); published in Q8_0 and converted as it is |
Qwen | 0.64 GB | 0.64 GB | beside the chat model |
Qwen3 has other dense models - 1.7B, 8B, 14B, 32B - which declare the same architecture as the 4B and so should convert and run the same way. None of them has been tried here, and until one is, that is a reading of the file and not a claim. (Qwen3-30B-A3B is not one of them: the number is close but it is a mixture of experts, with an architecture of its own.)
Any other Q4_K_M of the same models converts and runs just as well, including the "dynamic" ones that give different tensors different types: the engine reads Q4_K, Q5_K, Q6_K and Q8_0, and refuses a file carrying anything else rather than guessing. The four above are the files the numbers in this repository were measured on, and the ones MODELS.md gives fingerprints for, so that a converted file can be checked against the one measured here without downloading anything.
All four models are published under the Apache 2.0 licence by Alibaba Cloud, and quantized into GGUF by them or by the people above. Their licence is theirs, not this project's: MODELS.md says what that means if you pass a converted file on to somebody else.
Download. Nothing needs installing — one file, resumable:
mkdir -p models/gguf && cd models/gguf
curl -L -C - -O \
https://huggingface.co/bartowski/Qwen_Qwen3.6-35B-A3B-GGUF/resolve/main/Qwen_Qwen3.6-35B-A3B-Q4_K_M.gguf
cd ../..With pip install huggingface_hub you can instead fetch a whole quantization,
which is handier when it is split into several files, as Qwen3-Coder-Next is:
huggingface-cli download Qwen/Qwen3-Coder-Next-GGUF \
--include "Qwen3-Coder-Next-Q4_K_M/*" --local-dir models/ggufConvert. Python 3, no packages needed. With a model split into several files, give the first one and the rest is found:
python3 tools/gguf2jns.py models/gguf/Qwen_Qwen3.6-35B-A3B-Q4_K_M.gguf \
models/qwen3.6-35b-a3b-flat.jnsIt says what it is doing, and refuses a model it does not know with a clear message rather than half a conversion.
Cut the experts into bit planes. Worth it on any machine that cannot hold all the experts in memory, which is the case this engine is built for: the down matrix is stored as three planes of two bits, so the engine can read four bits of it, or two, instead of six, and decide which when the model is opened. Nothing is lost — all three planes give the weights the file came with, bit for bit. This is the form the models here are kept in:
bin/x86_64-linux/jns_planes models/qwen3.6-35b-a3b-flat.jns \
models/qwen3.6-35b-a3b.jns
rm models/qwen3.6-35b-a3b-flat.jnsAfterwards:
bin/x86_64-linux/jns_check models/qwen3.6-35b-a3b.jns # what is inside
bin/x86_64-linux/jns_check models/qwen3.6-35b-a3b.jns --verify # and the checksumsBoth steps are deterministic, so the file can be checked against the fingerprints in MODELS.md.
Room needed: three copies of the model at the worst moment — the GGUF, the converted file and the one with the planes — so about 145 GB for the largest. The flat file goes as soon as the planes are written, and the GGUF can go too, though keeping it saves the download the day a new version of the converter is worth running.
The multi-token prediction block (Qwen3-Next only, optional, faster replies). Its own model carries it, GGUF files leave it out, so it comes from the original checkpoint's last shard:
huggingface-cli download Qwen/Qwen3-Next-80B-A3B-Instruct \
--include "model-00041-of-00041.safetensors" --local-dir models/hf
bin/x86_64-linux/hf2jns_mtp models/qwen3-next.jns \
models/hf/model-00041-of-00041.safetensors models/qwen3-next-mtp.jnsAnd chat:
bin/x86_64-linux/janas-chat models/qwen3-next.jns \
--mtp models/qwen3-next-mtp.jns --statsThe first start reads the model's resident weights (a couple of GB) and then fills the expert cache in the background, while you read and type. The second reply of a session is the one that shows the machine's real speed.
INSTALL.md has the details, including the other distributions and what to do when the GPU is not found.
janas-chat <model.jns> [options]
| Option | What it does |
|---|---|
--mtp <file> |
the model's multi-token prediction block: faster replies, same text |
--draft <file> |
a small model of the same family to guess the next tokens: faster replies, same text (see below) |
--ctx <tokens> |
context length (default 16384, never more than the model was trained for) |
--cache <GiB> |
memory for streamed experts (default: what the machine can spare) |
--reserve <GiB> |
when the cache is automatic, the memory left to other programs (default: a fifth of the machine's); more keeps a busy desktop out of swap, at the price of a smaller cache |
--bits <2|4|6> |
bits per weight of the experts' down matrix (default: the memory decides) |
--attention <exact|fast> |
how the attention scores are computed (default: fast above 16384 tokens of context) |
--no-preload |
do not fill the expert cache at start from this machine's profile |
--no-recap |
when the context fills, drop the oldest turns without summing them up |
--system <text> |
the system message ("" for none; the default gives the assistant its name, Janas, names the model it thinks with, and says who wrote the engine) |
--temp, --top-k, --top-p, --min-p, --seed |
sampling |
--max <tokens> |
longest reply (default: no limit — a model can fill the whole context) |
--no-spec |
no speculative decoding (slower, and the same text: see below) |
--no-markdown |
print the model's marks instead of reading them |
--no-gpu |
never give the GPU work, whatever the mode says |
--mode auto|eco|max |
power mode: eco never uses the GPU for speed alone |
--think on|off |
reasoning before replying, for models that do it |
--stats |
a line of counters after every reply |
On --mtp, --draft and --no-spec. Speculative decoding is not a trade of
quality for speed: it changes how many tokens come out of one pass, never
which ones. The small multi-token prediction block guesses the next few
tokens, the model then runs one pass over the guesses, and a guess is kept
only where the token the model itself sampled is the very same one - at the
first disagreement the guess is thrown away and the model's own token stands.
So no wrong guess can survive into the text, and with the same seed the reply
is the one you would have got without any of it - checked on every run of the
tests, 256 tokens with drafts and 256 without, from one seed, identical. On
the development laptop --mtp is worth about a third on Qwen3-Next-80B-A3B,
which answers at 23 tok/s without it and 30 with.
--no-spec and /spec off turn it off; they buy nothing but time, and are
there to measure with.
Where a model has no prediction block the guesses are copied from the
conversation instead, which costs nothing and is right about a quarter of the
time - enough to be worth it on a dense model and barely enough on a mixture.
--draft puts a small model of the same family in their place: it has to
share the vocabulary, it borrows the big model's threads, and it predicts
where copying only repeats, so it is right about three quarters of the time.
Whether that pays depends on what a second token in a pass costs, which is
about 6% on a dense model and about 28% on a mixture, where two tokens route
to different experts. So it is for dense models: Qwen3-0.6B drafting Qwen3-4B
gives 1.14x - 23.0 to 26.3 tok/s, measured at the chat's own
temperature, which is why the pair is lower than the table above - and on
Qwen3-30B-A3B it would not pay for itself.
On sampling. The defaults are the ones Qwen recommends for its own models:
temperature 0.7, top-k 20, top-p 0.8. They are narrower than they look. With
top-p 0.8, on Qwen3-4B, 196 of 256 tokens had a single candidate left after
the cut - nothing to draw, so the reply is the one the model would have given
greedily, and two questions that mean the same thing get the same words.
Starting with --top-p 0.95 brings that to 142 of 256. It is not a fault and
not a setting this project chose: it is what those numbers do to a model that
is sure of itself.
On --attention. At a long context most of a token goes into attention,
and most of that into the scores: one dot product per position, per head. The
fast way reads the query as sixteen-bit integers with a scale of its own, so
the dot product becomes an exact integer sum — the machine does sixteen of
those products at a time instead of eight, and the scores come out a fifth to
a quarter quicker. What it costs is the query's quantization, a relative
3e-5, sixty times less than the eight-bit keys and values already carry; on
four hundred tokens against llama.cpp the most likely token agreed 394 times
out of 400 on C source against 393 the exact way, and 377 against 382 on
Italian prose. The engine asks for it by itself only when you asked for more
than 16384 tokens of context, where the gain is real and the conversation
will get there; --attention exact and --attention fast settle it either
way, and so does JANAS_ATTN=float|int16.
In the chat, a line starting with / is a command, and it turns grey-blue as
soon as it names one the chat really has:
| Command | |
|---|---|
/help |
the commands, on a page of its own |
/about |
what this is, who wrote it, and what it is running on |
/stats [on|off] |
the counters after each reply |
/reset |
a new conversation, and a clean window |
/system [text|off] |
show or change the system message |
/temp <t>, /spec, /think |
sampling, drafts, reasoning |
/mode [auto|eco|max] |
power mode, and what the engine chose |
/experts [n] |
experts per token: fewer is faster and a little less accurate |
/markdown [on|off] |
read the model's marks, or print them |
/gpu [on|off] |
whether the GPU may be given work |
/save-config, /load-config |
keep these settings for the next start, or read them back; both ask first where something would be lost |
/del-config |
forget the saved settings: the next start uses the chat's own |
/quit |
leave (also Ctrl-D) |
/help and /about open a page over the conversation: the text is laid out
to the window however narrow it is, the arrows and Page Up and Page
Down move in it, and ESC closes it - no other key does, so a page cannot
be shut by the one you pressed looking for a way down. /about gives the
version, the author, the licence and the repository, and then the model, the
machine and the engine as they are at that moment: the three things a report
of a problem needs, which were until now to be gathered by hand.
What a command changes lasts for that conversation and no longer: a knob
turned to try something does not quietly become the way things are.
/save-config writes the settings to ~/.config/janas/chat.conf and
/load-config reads them back, losing whatever is set at the time, which it
asks about first, and /del-config throws the file away so that the next
start uses the chat's own settings again. The file
holds the words of the command line, one option to a line and the rest of the
line its value, so the same reader takes both and you can open it and edit it;
it is read before the command line, which therefore wins. The system message
is not kept in it, and the context, the cache, the bits and the attention are
read at the next start, since they are settled when the model is opened.
A model marks its replies up - **bold**, *italic*, backticks, fenced
blocks, headings, bullets - and the chat reads those marks rather than
printing them: bold is bold, a block of code is told from the prose by its
colour, a dash becomes a bullet. An underscore counts as italic only alone and
at the edge of a word, so _this_ and **_this_** are marks while
snake_case and __init__, which a conversation about code is full of, stay
as they are; a star inside a word is not a mark either, so 2*3 is a product.
--no-markdown and /markdown off print the marks as they come, and a pipe
always gets them: a script may want them.
The box the message is written in takes the rows the message needs, up to a third of the window, and goes back to one when it has been sent: a long question is read while it is written instead of scrolling away sideways.
The conversation is kept as it is written, so Page Up and Page Down look back over it and Ctrl-Home and Ctrl-End go to its beginning and its end. The line being written stays where it is while you read, and sending brings the end back. A terminal cannot do this by itself here: the chat pins its footer with a scrolling region, and the lines that leave the top of one never reach the terminal's own history.
The arrows move in the message - up and down between its lines while it has
more than one - and bring back earlier prompts from its first and last line,
kept between sessions; Home, End, Delete, Ctrl-A/E/K/U/W and Ctrl-L do what they do
everywhere; Tab completes a command; a backslash and Enter add a line to
the message rather than sending it, the backslash itself is not part of what
the model reads, and the box grows to hold what is written; Ctrl-C stops a
reply.
Piped in or out, the chat writes plain text with no drawing at all, so it can be scripted:
printf 'Explain mixture of experts in one paragraph.\n/quit\n' \
| bin/x86_64-linux/janas-chat qwen3-next.jns --temp 0 --max 200The library is shaped for foreign-function interfaces from the start: opaque
handles, fixed-width types, nothing passed by value that is not a number,
parameter blocks that carry their own size, error codes as int32, and reply
text returned on request as whole UTF-8 characters. It is declared for C
and for FreeBASIC.
janas_llm *llm;
janas_llm_chat *chat;
struct janas_llm_params p;
janas_llm_params_default(&p);
janas_llm_open("qwen3-next.jns", &p, &llm);
janas_llm_chat_create(llm, NULL, &chat);
janas_llm_chat_send(chat, "Hello! Introduce yourself.", -1);
char piece[256];
int32_t len;
while (janas_llm_chat_next(chat, piece, sizeof(piece), &len) == JANAS_LLM_OK)
fwrite(piece, 1, (size_t)len, stdout);cc yours.c -Iinclude -Lbin/x86_64-linux -ljanas_llm -o yoursA program that keeps the conversation itself - an HTTP client does, and sends
it whole every time - hands it over with janas_llm_chat_load: the messages
with their roles, the last one the user's. Whatever the new conversation
shares with the one already computed is not read again, so the same messages
plus a new one cost only the new one. janas_llm_chat_prompt continues raw
text, with no chat format around it, and janas_llm_default_system gives the
system message Janas's own programs use. janas_llm_chat_stats says why a
reply ended and how much of its prompt was already computed.
janas_llm_chat_keep keeps a few conversations computed besides the current
one, for a program that switches between them: the one that shares most of
the next request is copied back instead of being read again.
include/janas/llm.bi declares the same library for
FreeBASIC, and so for BASIC MODERN, the dialect implemented by Prabanta —
the platform Janas is meant to sit inside. The same conversation as above:
#include once "janas/llm.bi"
dim as janas_llm ptr llm
dim as janas_llm_chat ptr chat
dim as janas_llm_params p
janas_llm_params_default(@p)
janas_llm_open("qwen3-next.jns", @p, @llm)
janas_llm_chat_create(llm, NULL, @chat)
janas_llm_chat_send(chat, "Hello! Introduce yourself.", -1)
dim as zstring * 256 piece
dim as long n
do
'' the piece is not closed by a NUL: one byte is left for it
dim as long r = janas_llm_chat_next(chat, @piece, sizeof(piece) - 1, @n)
if r <> JANAS_LLM_OK then exit do
piece[n] = 0
print piece;
looptools/chat.bas is a whole chat in a hundred lines, sampling
and counters included:
fbc tools/chat.bas -i include -p bin/x86_64-linux -x janas-chat-fb
LD_LIBRARY_PATH=bin/x86_64-linux ./janas-chat-fb qwen3-next.jns
echo "Hello! Introduce yourself." | LD_LIBRARY_PATH=bin/x86_64-linux ./janas-chat-fb qwen3-next.jnsTwo things to know. FreeBASIC does not tell names apart by their case, so the
constant the C header calls JANAS_LLM_ABI_VERSION is JANAS_LLM_ABI there —
it would otherwise collide with the function janas_llm_abi_version(). And its
line input reads the terminal, not the standard input: to work in a pipe too,
chat.bas asks isatty and reads a pipe through the CONS device.
bin/x86_64-linux/janas-server qwen3-next.jnsjanas-server puts the model behind an HTTP API that follows
OpenAI's, on
http://127.0.0.1:8080/v1, so that a client written for that API can use a
model running on this machine. It takes the model options janas-chat takes
(--ctx, --cache, --reserve, --mtp, --draft, --mode, --no-gpu) and these:
| Option | |
|---|---|
--host ADDR |
the address to listen on; 127.0.0.1 (the default) is this machine only, 0.0.0.0 opens it to the network |
--port N |
8080 by default |
--api-key KEY |
ask every client for this key, as Authorization: Bearer KEY; JANAS_API_KEY sets it too |
--name ID |
the model's name for clients; the file's name without .jns by default |
--think on|off |
reasoning before replying, for models that do it, when a request does not say |
--test |
serve a chat page for trying it out, on /test (below) |
--keep N |
conversations kept computed besides the one in use (8; 0: none), so that clients taking turns, or a client's requests on the side for titles and tags, do not have their conversations read again from the start |
--keep-disk GIB |
disk for them below the memory, in ~/.cache/janas, so that they outlive the server: a long system message and tools are read once (8; 0: none; never more than a quarter of the space left). Written only when a conversation leaves the memory and when the server stops |
--keep-memory GIB |
the memory for them; by default half of what the expert cache and the margin left to other programs leave free at start, at least 256 MiB |
--store DIR |
where stored completions, responses and conversations are kept, so that they outlive the server (~/.local/share/janas/server) |
--no-store |
keep them in memory only |
--queue N |
requests allowed to wait (16) |
--max-body MIB |
the largest request body (32) |
--connections N |
open connections at most (64) |
--verbose |
a line per request on the standard error |
What it answers:
| Request | |
|---|---|
GET /v1/models, GET /v1/models/{model} |
the models it runs, with a description of how they run here; DELETE answers that a model read from a file is not deleted through the API |
POST /v1/chat/completions |
a conversation, the reply whole or streamed ("stream": true, server-sent events), with tools, JSON output, log-probabilities and more than one reply (below) |
GET, POST, DELETE /v1/chat/completions/{id}, GET /v1/chat/completions, .../messages |
the completions a client asked to keep ("store": true): read, listed with their metadata, changed, deleted |
POST /v1/completions |
raw text to continue, no chat format around it; several prompts, token ids, echo, logprobs, best_of, suffix |
POST /v1/responses and the six operations under it |
OpenAI's newer API: items in and out, previous_response_id, conversations, streamed events, background and cancel, input_items, input_tokens, compact |
/v1/conversations and its seven operations |
conversations the server keeps, and their items |
POST /v1/embeddings |
with --embedding-model FILE: the vectors of texts (below) |
POST /v1/moderations |
asked of the chat model (below) |
GET /health |
{"status":"ok"} |
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"messages": [{"role": "user", "content": "What is the capital of Sardinia?"}]}'A client that lets you set the address of an OpenAI-compatible API should
need nothing else: the base URL is http://127.0.0.1:8080/v1, and the key
whatever you gave --api-key (anything, if you gave none). It has been tried
here with curl and with its own test page, not yet with particular clients.
The conversation is the client's. OpenAI's API keeps no state: every
request carries the whole conversation. janas-server keeps the one it
computed last, and reads again only what a request adds to it: a client that
sends the same messages plus a new one pays for the new one alone, and
usage.prompt_tokens_details.cached_tokens says how much was already
computed. The replies it wrote are remembered with their reasoning, so the
answer a client sends back - without the reasoning, as clients do - still
counts as the same turn. A conversation that differs earlier is read again
from where it differs, and from the start on models with a recurrent state
(Qwen3-Next, Qwen3.5 and 3.6).
Reasoning comes apart from the answer, in reasoning_content, whole or
streamed. "reasoning_effort": "none" turns it off for a request, any other
effort turns it on, and "chat_template_kwargs": {"enable_thinking": false}
works too; --think decides for requests that say nothing.
Beyond OpenAI's fields, which clients ignore: top_k and min_p in a
request, and in every reply's usage a janas object with the context used
and its size and the seconds spent reading the prompt and writing the reply.
Tools. tools, tool_choice (none, auto, required, a function)
and parallel_tool_calls, and the older functions. The tools are written
into the system message the way the model's chat template writes them, and a
call the model opens is held to a grammar of the functions and their
parameters until it closes it, so its arguments are always a JSON object
valid against the function's schema. The calls come back in tool_calls,
whole in one delta of the stream, and the reply ends with
"finish_reason": "tool_calls"; sent back with the tools' answers, they are
found again among the replies the server remembers, so the next turn reads
only the answers. Qwen models write calls in one of two ways, and the
template says which: as JSON (Qwen3, Qwen3-Next) or as XML (Qwen3.5 and 3.6,
Qwen3-Coder). Both have been run here: the JSON way on Qwen3-4B and
Qwen3-0.6B, the XML way on Qwen3.6-35B-A3B (two calls at once, their answers
sent back, 422 of 471 prompt tokens reused on the next turn). Qwen3-Coder's
own way of describing the tools is written from its template and not yet
run.
JSON output. response_format json_object or json_schema (and
text.format in the Responses API): the answer is held to a grammar token by
token, so it is valid JSON when the model ends it (not when max_tokens cuts
it). A schema is held to its types, the properties of an object in the
schema's order with the required ones always there and no others, array
items and bounds up to 32, string lengths up to 64, enum, const,
anyOf, oneOf, $ref within the schema and nullable; pattern,
format and numeric bounds are not enforced. The reasoning, where there is
one, stays free.
Log-probabilities. logprobs and top_logprobs (up to 20) in chat,
logprobs in completions, top_logprobs in responses. They are the
model's own: where a grammar forced a token the model found unlikely, its
log-probability says so.
More than one reply with n (and best_of in completions, which keeps
the best by mean log-probability); presence_penalty, frequency_penalty
and logit_bias as OpenAI defines them. The replies of n are written one
after the other on the same prompt.
Embeddings come from a second, small model opened with
--embedding-model, such as Qwen3-Embedding-0.6B: texts or token ids, one or
many, dimensions to cut the vector shorter, floats or base64. Against
llama.cpp on the same file the vectors of the same text agree to a cosine of
0.98 to 0.996, and the similarities between texts come out alike.
Moderations are asked of the chat model: at temperature 0 it answers
OpenAI's thirteen categories with a JSON object held to a schema, and each
score is the chance it gave true against false in that place. It is a
general model following instructions, not a classifier trained for the task.
Kept on disk. Stored completions, responses and conversations are
written, one file each, to ~/.local/share/janas/server (or --store DIR),
and read back when the server starts again; the oldest go past a thousand of
each. The directory and its files are readable by their owner alone, since
they hold what the clients said, and one server at a time uses a directory.
--no-store keeps them in memory only, until the server stops. A response
running in the background when the server is killed is found as failed at
the next start. Images and sound in messages are refused with a 400 that
says so.
Every other operation of OpenAI's API - there are 345 in the version it
follows (2.3.0) - has a route that answers 501 and says why: not written
yet, needs a model of another kind (sound, images, video), trains models, or
belongs to the administration of OpenAI's service. The table is generated
from the specification by tools/openapi_routes.py, which also counts what is
done.
One request at a time. The model holds one sequence, so requests wait in
line and run in the order they came; when the line is full the answer is 429.
While a streamed request waits, the stream says where it stands in the line
(: queue N, a comment clients skip) and the test page shows it.
A client that closes the connection while a reply streams stops the reply.
On the network. It listens on this machine alone unless --host says
otherwise, and then it warns if there is no key. There is no HTTPS yet: across
a network you do not trust, put it behind something that adds it.
The test page. With --test, http://127.0.0.1:8080/test is
Janas-Chat Web: a chat in the browser that talks to the API like any client,
with the reasoning in grey, a line under each reply with the context used and
the speed (last, lowest, highest and mean of the conversation), and the
settings - system message, temperature, reasoning, key - behind Settings.
It reads the marks a model writes as janas-chat does, and also quotes,
tables, struck text and links; a link asks before it opens, because links an
AI writes can be wrong or unsafe. Markdown sample shows a text of ours with
every mark it knows, without asking the model. The page is inside the
program, loads nothing from anywhere else and may talk only to this server.
bin/x86_64-linux/janas-bench qwen3-next.jnsClose what you can before running it: it measures the machine you give it, and
a browser with thirty tabs is part of the machine. It tries the configurations
the machine offers — how many threads, which cores,
with and without the GPU — on real passes of the model, keeps the fastest,
writes them into ~/.cache/janas so the next start begins from them, and leaves
a report in the current directory. On the development laptop that report says,
among other things, that bringing the efficiency cores in makes decoding about
a tenth faster than the performance cores alone - and that for blocks of
several tokens, which is what a prompt and a draft check are, the engine often
wants fewer threads than for one.
tools/gguf2jns.py |
GGUF → the .jns format the engine reads |
tools/hf2jns_mtp |
the multi-token prediction block, from the original checkpoint |
tools/jns_planes |
rewrites a model with the experts' down matrix in bit planes, so a machine short of memory can read part of it |
jns_check |
checks a model file, with --verify every expert's checksum |
tools/openapi_routes.py |
janas-server's route table, from OpenAI's specification |
llm_eval |
compares the engine's logits against a reference dump |
bench_gemm, bench_attn, bench_long, ... |
the pieces measured on their own |
They are built into bin/x86_64-linux/ along with everything else.
This is where help is worth most. The engine decides everything from what it measures, but it has only ever measured one machine, so the decisions are tuned to one shape of CPU, one disk and one GPU.
What to send. Run the benchmark and attach the report it writes (a
janas-bench-<date>.txt in the current directory):
bin/x86_64-linux/janas-bench <model.jns>The report already carries the CPU's name, the core layout, the GPU if there is one, and the speed of every configuration tried. Add:
- the distribution and kernel (
uname -a), and the RAM (free -g); - the disk the model sits on (NVMe? SATA? over USB?);
- the model you used and its file size;
- what you saw: numbers that look wrong, a reply that made no sense, a crash, a machine that went unresponsive — with what you were doing at the time.
The hardware report form asks for exactly that, in that order. There are two more forms, for something that went wrong and for a model that will not convert.
What is most wanted, in order.
- CPUs without AVX-VNNI, and AMD — the arithmetic has a path for them and it is the least exercised of all.
- Discrete GPUs (NVIDIA, AMD, Intel Arc). The GPU support is written and
correct, bit for bit, but on an integrated GPU it gains nothing: whether it
pays on a real card is unknown. Careful: a GPU driver reset takes the
desktop with it. Start with
janas-bench, not with a long chat. - Machines with 16 GB or less. The engine is supposed to trade quality for memory on its own — fewer bits per weight, fewer experts per token — and the thresholds for that were measured on 32 GB.
- Slower disks. Everything assumes an NVMe SSD; on a SATA disk the engine should still work and simply wait more, but nobody has watched it do so.
- Other models of the supported families. A file that will not convert, or converts and then answers nonsense, is a useful bug.
What not to bother with yet: Windows and macOS (not supported), and models outside the Qwen3 MoE families (they are refused with a message).
Patches are welcome too, but a measurement from a machine nobody here can buy is worth more than most patches.
Janas is free software under the GNU General Public License, version 3 or later — see LICENSE. Every source file carries its SPDX line.
janas-server, and only it, has two libraries of others compiled in, their
sources unchanged in src/third_party/ with their
licences and where they came from: GNU libmicrohttpd 1.0.10 (LGPL 2.1 or
later, or the eCos licence without HTTPS, as it is built here) and yyjson
0.13.0 (MIT). libjanas_llm has none.
The models are not covered by it. A converted .jns file is a derivative of
the model, not of Janas, and keeps the licence its authors gave it;
MODELS.md says what this project will and will not redistribute,
and what to check before you pass a converted model on.
CONTRIBUTING.md. The short version: a measurement from a
machine nobody here owns is worth more than most patches, and every commit needs
a Signed-off-by line (git commit -s).