-
Notifications
You must be signed in to change notification settings - Fork 2
Installation and Configuration
How TinyTitan is installed, where it puts things, which model to install, and every setting the tools and the engine read. For the generation and engine parameters themselves, see Runtime Controls; for the API, Local Server and API.
Two paths, one result: built arm64 binaries, the tools, and a tinytitan
command that starts the server.
One command — checks the Mac, downloads the newest published release, verifies its checksum, offers a model, and installs the launcher:
bash -c "$(curl -fsSL https://raw.githubusercontent.com/Pummelchen/TinyTitan/main/tools/install_tinytitan.sh)"Use that form, not curl … | bash: a pipe makes the script's stdin the pipe,
so it cannot ask anything and silently takes the default at every step —
including "no model". From a clone, bash tools/install_tinytitan.sh does the
same and never needs chmod +x.
| Flag | Does |
|---|---|
--yes, -y
|
Answer yes to every question (unattended install) |
--model NAME |
Install this model without asking; through a pipe the default is ornith15-8bit
|
--no-model |
Install no model |
--web |
Also set up the browser chat window |
--no-web |
Do not offer the browser chat window |
--version TAG |
Install that release instead of the newest (e.g. --version v5.12) |
--from-source |
Clone and swift build instead of downloading a release (needs Xcode) |
--dir PATH |
Where --from-source clones to (default ~/TinyTitan) |
It needs no Homebrew, Xcode, git, Node or Python, and it never deletes a model.
Everything it creates lives under ~/.tinytitan and ~/.local/bin, so removing
those two directories removes the install. Re-running it is safe: it replaces the
engine and tools in place and keeps the models.
From source — for contributors and for machines you build yourself:
git clone https://github.com/Pummelchen/TinyTitan.git
cd TinyTitan
swift build -c releaseNeeds Swift 6.4+ (Xcode 27 or a matching toolchain) on an arm64 Mac with
macOS 26 or later. Products land in .build/release; add
--product TinyTitanServer to build only the server.
tools/install_models.sh # what is installed, what staging is kept
tools/install_models.sh --choose # the install menu
tools/install_models.sh qwen35-4b # one width
tools/install_models.sh katcoder both # 4-bit and 8-bit from ONE download
tools/install_models.sh --all-4bit # every 4-bit model
tools/install_models.sh clean # drop staging that cannot be reused
tools/install_models.sh --help # every key, its source and disk size| Model | Keys | Notes |
|---|---|---|
| Ornith 1.5 35B-A3B |
ornith15, ornith15-8bit, ornith15-mtp
|
The default install is ornith15-8bit
|
| Qwen 3.6 35B-A3B |
qwen36, qwen36-8bit, qwen36-mtp
|
|
| Qwen3.8-Flash-Next 125B-A6B |
qwen38flash, qwen38flash-8bit, qwen38flash-mtp
|
One ~360 GB fetch per width |
| Qwen-AgentWorld 35B-A3B |
agentworld, agentworld-8bit
|
Tuned for agentic tool use |
| KAT-Coder-V2.5-Dev 35B-A3B |
katcoder, katcoder-8bit
|
Kwaipilot's Qwen 3.6 fine-tune for coding |
| Qwen 3.5 2B / 4B / 9B |
qwen35-2b, qwen35-4b, qwen35-9b (+ -8bit) |
Dense; the only models the CPU engine runs |
Use both when you want two widths of the same model. The Qwen3.5-MoE
checkpoints convert their 4-bit and 8-bit builds in a single pass from one
~70 GB fetch, so asking for one width at a time fetches the checkpoint twice.
Installing a single width still converts both and keeps the other snapshot, so
adding the second width later reuses it.
Where it goes. Models install to $TINYTITAN_MODELS_DIR — the installer
sets ~/.tinytitan/models, a checkout uses <repo>/models — one directory per
model with its verified-install.json receipt beside it. A download and its
conversion are staged under $TINYTITAN_WORK_DIR (default <repo>/.build).
Staging is kept exactly while it can save a fetch: the converted snapshot lives
until the other width of that model is installed.
Disk checks come first. A download is refused before it starts when either
volume cannot hold it, naming both numbers and both ways out (free space, or
TINYTITAN_WORK_DIR on another volume). TINYTITAN_SKIP_DISK_CHECK=1 starts
anyway. A resumed conversion does not need the room again. An interrupted
download continues with --resume:
swift run -c release TinyTitanRepack \
--model ornith15-8bit --output models/ornith-1.5_35B_A3B_8Bit --resumeTo fetch through a Hugging Face mirror, set the Hub's own variable (the
converter also takes --endpoint): HF_ENDPOINT=https://hf-mirror.com tools/install_models.sh qwen38flash. Set HF_TOKEN only if Hugging Face
requests authentication.
Every install is built from the model's own bf16 release and quantized here (group-64 affine); third-party quantizations are deliberately not used.
- Update the engine: re-run the installer. It replaces the binaries and tools in place and keeps every model.
-
Update from source:
git pull && swift build -c release. -
Remove a model: delete its directory under
models/. Nothing else refers to it, andtools/install_models.shstops listing it. -
Remove everything: delete
~/.tinytitanand~/.local/bin/tinytitan*for an installer-created setup, or the checkout andmodels/for a source build.
tools/server_launcher.sh starts the server, or starts it and wires one coding
client to it. Interactive by default; every question has a default.
tools/server_launcher.sh # ask me everything
tools/server_launcher.sh --client server --model qwen35-4b --bits 4
tools/server_launcher.sh --client codex --model ornith --bits 8 --ram 8
tools/server_launcher.sh --dry-run --model qwen38 --bits 4 # print, start nothing| Flag | Values | Default | Meaning |
|---|---|---|---|
--client |
server, codex, claude, qwen, opencode, zed
|
server |
Start the API alone, or the API plus that client's config |
--model |
install key or catalogue id | asked | Which model to serve |
--bits |
4, 8
|
asked | Width for a model key |
--engine |
cpu, gpu
|
the install's own | Which engine serves it; dense Qwen 3.5 is the only real choice |
--mode |
fast, full
|
full |
fast strips coding-CLI boilerplate from the prompt |
--answers |
default, concise
|
default |
Concise adds the shorter-answer system prompt |
--thinking |
off, on, or a level the model lists |
off |
Reasoning mode |
--ram |
whole GB ≥ 4 | install profile | Target for the whole server process (GPU models) |
--context |
number, native, max
|
native (262144) |
Prompt plus response capacity; --yarn enables 524288/1048576 |
--kv |
4, 8, 16
|
8 |
KV-cache storage precision |
--yarn |
flag | off | Enable YaRN context scaling |
--port |
1–65535 | 8080 |
Server port (TINYTITAN_PORT overrides) |
--concurrency |
power of two ≤ 256 | 1 |
Generations served at once; buys fairness, not throughput, and turns the prompt cache off |
--memory |
flag | off | Enable persistent agent memory for this working directory |
--web |
flag | off | Also open TinyTitan's browser chat window on port 7788 |
--dry-run |
flag | off | Print the server command and client setup; start nothing |
--prompt-cache |
multi-prefix, off
|
multi-prefix |
Prompt-state reuse (256 MiB) |
--mtp-model |
directory | off | Attach a native speculative draft head (GPU only) |
--mtp-memory-mib |
MiB | 384 |
Budget for that draft head |
Positionally the same launcher accepts
[<client>] [fast|full] [<model> [4|8]] [default|concise] [<thinking>] [<ram>],
so older command lines keep working.
| Variable | Effect | Default |
|---|---|---|
TINYTITAN_MODELS_DIR |
Where installs live |
models/ in a checkout, ~/.tinytitan/models for an installed copy |
TINYTITAN_WORK_DIR |
Where a download and its conversion are staged | <repo>/.build |
TINYTITAN_BIN_DIR |
Where the tools look for the release binaries | .build/release |
TINYTITAN_SKIP_DISK_CHECK |
1 starts a download the free-space check would refuse |
off |
HF_ENDPOINT |
Download through a Hugging Face mirror | unset |
HF_TOKEN |
Authenticate to Hugging Face when it asks | unset |
TINYTITAN_PORT |
Server port the launcher and route use | 8080 |
TINYTITAN_THINKING_MODE |
Default reasoning mode | off |
TINYTITAN_MODELS_DIR + TINYTITAN_CATALOG_JSON
|
Serve a catalogue captured earlier instead of asking the server for it | unset |
TINYTITAN_LAUNCHER_DRY_RUN |
1 prints the server command and starts nothing |
off |
TINYTITAN_WARM, TINYTITAN_WARM_TIMEOUT, TINYTITAN_WARM_TOKENS
|
Warm-up after launch: on/off, its curl timeout (s), its prompt size |
1, 900, 4000
|
TINYTITAN_ROOT |
Install root for the one-command installer | ~/.tinytitan |
| Variable | Effect | Default |
|---|---|---|
TINYTITAN_MEMORY |
1/on/true turns the subsystem on |
off |
TINYTITAN_MEMORY_DIR |
Where the store lives |
~/.tinytitan/memory for an install, memory/ in a checkout |
TINYTITAN_MEMORY_TOOLS |
full exposes the six memory_* functions to the model; off hides them |
configuration default |
TINYTITAN_MEMORY_NAMESPACE, TINYTITAN_MEMORY_USER, TINYTITAN_MEMORY_WORKSPACE
|
Namespace, user and workspace the store is scoped by |
tinytitan, the OS user, derived from TINYTITAN_WORKSPACE_DIR
|
TINYTITAN_WORKSPACE_DIR |
Launch directory; the launcher sets it to $PWD, which is what scopes memory per project |
unset for a bare server |
TINYTITAN_MEMORY_CACHE_MIB |
Cap on the in-memory cache | unset |
TINYTITAN_MEMORY_RETENTION_DAYS, TINYTITAN_MEMORY_MAX_WORKSPACES
|
Retention window and workspace cap | unset |
The subsystem runs inside the server process: there is no database to install
and nothing to start. See docs/agent-memory.md in the repository for the tool
list and the store format.
| Variable | Effect | Default |
|---|---|---|
TINYTITAN_DSH_ROOT |
Private root for TinyTitan's own DeepSeek Harness | ~/.tinytitan/dsh |
TINYTITAN_DSH_PORT |
Its browser UI port |
7788 (next free if taken) |
TINYTITAN_DSH_VERSION, TINYTITAN_DSH_NODE_VERSION, TINYTITAN_DSH_PNPM_VERSION
|
The pinned harness, Node and pnpm versions | pinned in the script |
TINYTITAN_DSH_DRY_RUN |
1 prints what would happen and changes nothing |
off |
TINYTITAN_DSH_NO_OPEN |
1 never opens a browser window |
off |
These change no result; they make the engine report what it did. They are the first thing to reach for when something is slow or wrong.
| Variable | Effect |
|---|---|
TINYTITAN_RUNNER_STATS |
Print the per-stage timing footer (prompt, embed, I/O, decode) |
TINYTITAN_KERNEL_STATS |
Per-command-buffer GPU timings |
TINYTITAN_KERNEL_SPLIT |
1 gives every kernel in the decode chain its own command buffer |
TINYTITAN_DECODE_IO_TRACE |
1 reports the decode phase's own expert I/O |
TINYTITAN_ROUTE_TRACE |
Names a file for the per-token routing trace (JSONL) |
TINYTITAN_PREFETCH_TRACE |
Names a file for the predictive-prefetch trace |
TINYTITAN_ACT_DUMP |
Names a directory to dump activations into |
TINYTITAN_MTP_TRACE |
1 traces each MTP draft/verify step on stderr |
TINYTITAN_ANE_MEMORY_TRACE |
1 traces the ANE runtime's arena footprint |
TINYTITAN_LOGIT_TRACE |
1 prints the top-2 logits per decode step |
TINYTITAN_WATCHDOGS |
1 enables the server's stall watchdogs |
-
tools/dsh_route.sh— print (or--write) the DeepSeek Harness route for the local server, generated from the installs you actually have:tools/dsh_route.sh,--write,--models qwen38 4,--reasoning off,--port,--context,--max-tokens,--provider,--from-server,--settings. Exits 2 when there is nothing it can describe. -
tools/dsh_local.sh— TinyTitan's own isolated DeepSeek Harness, the optional browser window:ensure,web,smoke,port,status,paths. It keeps its own home, npm prefix, pnpm store and port, and never reads or writes a DeepSeek Harness you run yourself. See TinyTitan Plugin. -
ttlanmanager— the LAN control plane for several Macs running DeepSeek Harness:top,list,prompt,prompt-all,workspace create,session archive,workspace delete, plus--peer HOST[:PORT],--key,--json,--timeoutand--once/--width/--height/--fromfor rendering an inventory without a fleet. See LAN Manager. -
TinyTitanRepack— the installer/converter behind the catalogue:--model <key> --output <dir>(with--resume,--overwrite),--input-snapshot <dir> --model-id <id>,--discard-partial --output <dir>, and--verify-install --input-gturbo <dir>to re-issue a receipt after a move.
Start
Use TinyTitan
DeepSeek Harness
Reference
Project