Skip to content

Installation and Configuration

André Borchert edited this page Sep 29, 2026 · 1 revision
TinyTitan

Installation and Configuration

How TinyTitan is installed, where it puts things, which model to install, and every setting the tools and the engine read. For the generation and engine parameters themselves, see Runtime Controls; for the API, Local Server and API.

Install the engine

Two paths, one result: built arm64 binaries, the tools, and a tinytitan command that starts the server.

One command — checks the Mac, downloads the newest published release, verifies its checksum, offers a model, and installs the launcher:

bash -c "$(curl -fsSL https://raw.githubusercontent.com/Pummelchen/TinyTitan/main/tools/install_tinytitan.sh)"

Use that form, not curl … | bash: a pipe makes the script's stdin the pipe, so it cannot ask anything and silently takes the default at every step — including "no model". From a clone, bash tools/install_tinytitan.sh does the same and never needs chmod +x.

Flag Does
--yes, -y Answer yes to every question (unattended install)
--model NAME Install this model without asking; through a pipe the default is ornith15-8bit
--no-model Install no model
--web Also set up the browser chat window
--no-web Do not offer the browser chat window
--version TAG Install that release instead of the newest (e.g. --version v5.12)
--from-source Clone and swift build instead of downloading a release (needs Xcode)
--dir PATH Where --from-source clones to (default ~/TinyTitan)

It needs no Homebrew, Xcode, git, Node or Python, and it never deletes a model. Everything it creates lives under ~/.tinytitan and ~/.local/bin, so removing those two directories removes the install. Re-running it is safe: it replaces the engine and tools in place and keeps the models.

From source — for contributors and for machines you build yourself:

git clone https://github.com/Pummelchen/TinyTitan.git
cd TinyTitan
swift build -c release

Needs Swift 6.4+ (Xcode 27 or a matching toolchain) on an arm64 Mac with macOS 26 or later. Products land in .build/release; add --product TinyTitanServer to build only the server.

Install a model

tools/install_models.sh                     # what is installed, what staging is kept
tools/install_models.sh --choose            # the install menu
tools/install_models.sh qwen35-4b           # one width
tools/install_models.sh katcoder both       # 4-bit and 8-bit from ONE download
tools/install_models.sh --all-4bit          # every 4-bit model
tools/install_models.sh clean               # drop staging that cannot be reused
tools/install_models.sh --help              # every key, its source and disk size
Model Keys Notes
Ornith 1.5 35B-A3B ornith15, ornith15-8bit, ornith15-mtp The default install is ornith15-8bit
Qwen 3.6 35B-A3B qwen36, qwen36-8bit, qwen36-mtp
Qwen3.8-Flash-Next 125B-A6B qwen38flash, qwen38flash-8bit, qwen38flash-mtp One ~360 GB fetch per width
Qwen-AgentWorld 35B-A3B agentworld, agentworld-8bit Tuned for agentic tool use
KAT-Coder-V2.5-Dev 35B-A3B katcoder, katcoder-8bit Kwaipilot's Qwen 3.6 fine-tune for coding
Qwen 3.5 2B / 4B / 9B qwen35-2b, qwen35-4b, qwen35-9b (+ -8bit) Dense; the only models the CPU engine runs

Use both when you want two widths of the same model. The Qwen3.5-MoE checkpoints convert their 4-bit and 8-bit builds in a single pass from one ~70 GB fetch, so asking for one width at a time fetches the checkpoint twice. Installing a single width still converts both and keeps the other snapshot, so adding the second width later reuses it.

Where it goes. Models install to $TINYTITAN_MODELS_DIR — the installer sets ~/.tinytitan/models, a checkout uses <repo>/models — one directory per model with its verified-install.json receipt beside it. A download and its conversion are staged under $TINYTITAN_WORK_DIR (default <repo>/.build). Staging is kept exactly while it can save a fetch: the converted snapshot lives until the other width of that model is installed.

Disk checks come first. A download is refused before it starts when either volume cannot hold it, naming both numbers and both ways out (free space, or TINYTITAN_WORK_DIR on another volume). TINYTITAN_SKIP_DISK_CHECK=1 starts anyway. A resumed conversion does not need the room again. An interrupted download continues with --resume:

swift run -c release TinyTitanRepack \
  --model ornith15-8bit --output models/ornith-1.5_35B_A3B_8Bit --resume

To fetch through a Hugging Face mirror, set the Hub's own variable (the converter also takes --endpoint): HF_ENDPOINT=https://hf-mirror.com tools/install_models.sh qwen38flash. Set HF_TOKEN only if Hugging Face requests authentication.

Every install is built from the model's own bf16 release and quantized here (group-64 affine); third-party quantizations are deliberately not used.

Update or remove

  • Update the engine: re-run the installer. It replaces the binaries and tools in place and keeps every model.
  • Update from source: git pull && swift build -c release.
  • Remove a model: delete its directory under models/. Nothing else refers to it, and tools/install_models.sh stops listing it.
  • Remove everything: delete ~/.tinytitan and ~/.local/bin/tinytitan* for an installer-created setup, or the checkout and models/ for a source build.

The launcher

tools/server_launcher.sh starts the server, or starts it and wires one coding client to it. Interactive by default; every question has a default.

tools/server_launcher.sh                                       # ask me everything
tools/server_launcher.sh --client server --model qwen35-4b --bits 4
tools/server_launcher.sh --client codex --model ornith --bits 8 --ram 8
tools/server_launcher.sh --dry-run --model qwen38 --bits 4     # print, start nothing
Flag Values Default Meaning
--client server, codex, claude, qwen, opencode, zed server Start the API alone, or the API plus that client's config
--model install key or catalogue id asked Which model to serve
--bits 4, 8 asked Width for a model key
--engine cpu, gpu the install's own Which engine serves it; dense Qwen 3.5 is the only real choice
--mode fast, full full fast strips coding-CLI boilerplate from the prompt
--answers default, concise default Concise adds the shorter-answer system prompt
--thinking off, on, or a level the model lists off Reasoning mode
--ram whole GB ≥ 4 install profile Target for the whole server process (GPU models)
--context number, native, max native (262144) Prompt plus response capacity; --yarn enables 524288/1048576
--kv 4, 8, 16 8 KV-cache storage precision
--yarn flag off Enable YaRN context scaling
--port 1–65535 8080 Server port (TINYTITAN_PORT overrides)
--concurrency power of two ≤ 256 1 Generations served at once; buys fairness, not throughput, and turns the prompt cache off
--memory flag off Enable persistent agent memory for this working directory
--web flag off Also open TinyTitan's browser chat window on port 7788
--dry-run flag off Print the server command and client setup; start nothing
--prompt-cache multi-prefix, off multi-prefix Prompt-state reuse (256 MiB)
--mtp-model directory off Attach a native speculative draft head (GPU only)
--mtp-memory-mib MiB 384 Budget for that draft head

Positionally the same launcher accepts [<client>] [fast|full] [<model> [4|8]] [default|concise] [<thinking>] [<ram>], so older command lines keep working.

Environment variables

Install, storage and the launcher

Variable Effect Default
TINYTITAN_MODELS_DIR Where installs live models/ in a checkout, ~/.tinytitan/models for an installed copy
TINYTITAN_WORK_DIR Where a download and its conversion are staged <repo>/.build
TINYTITAN_BIN_DIR Where the tools look for the release binaries .build/release
TINYTITAN_SKIP_DISK_CHECK 1 starts a download the free-space check would refuse off
HF_ENDPOINT Download through a Hugging Face mirror unset
HF_TOKEN Authenticate to Hugging Face when it asks unset
TINYTITAN_PORT Server port the launcher and route use 8080
TINYTITAN_THINKING_MODE Default reasoning mode off
TINYTITAN_MODELS_DIR + TINYTITAN_CATALOG_JSON Serve a catalogue captured earlier instead of asking the server for it unset
TINYTITAN_LAUNCHER_DRY_RUN 1 prints the server command and starts nothing off
TINYTITAN_WARM, TINYTITAN_WARM_TIMEOUT, TINYTITAN_WARM_TOKENS Warm-up after launch: on/off, its curl timeout (s), its prompt size 1, 900, 4000
TINYTITAN_ROOT Install root for the one-command installer ~/.tinytitan

Agent memory

Variable Effect Default
TINYTITAN_MEMORY 1/on/true turns the subsystem on off
TINYTITAN_MEMORY_DIR Where the store lives ~/.tinytitan/memory for an install, memory/ in a checkout
TINYTITAN_MEMORY_TOOLS full exposes the six memory_* functions to the model; off hides them configuration default
TINYTITAN_MEMORY_NAMESPACE, TINYTITAN_MEMORY_USER, TINYTITAN_MEMORY_WORKSPACE Namespace, user and workspace the store is scoped by tinytitan, the OS user, derived from TINYTITAN_WORKSPACE_DIR
TINYTITAN_WORKSPACE_DIR Launch directory; the launcher sets it to $PWD, which is what scopes memory per project unset for a bare server
TINYTITAN_MEMORY_CACHE_MIB Cap on the in-memory cache unset
TINYTITAN_MEMORY_RETENTION_DAYS, TINYTITAN_MEMORY_MAX_WORKSPACES Retention window and workspace cap unset

The subsystem runs inside the server process: there is no database to install and nothing to start. See docs/agent-memory.md in the repository for the tool list and the store format.

The browser chat window

Variable Effect Default
TINYTITAN_DSH_ROOT Private root for TinyTitan's own DeepSeek Harness ~/.tinytitan/dsh
TINYTITAN_DSH_PORT Its browser UI port 7788 (next free if taken)
TINYTITAN_DSH_VERSION, TINYTITAN_DSH_NODE_VERSION, TINYTITAN_DSH_PNPM_VERSION The pinned harness, Node and pnpm versions pinned in the script
TINYTITAN_DSH_DRY_RUN 1 prints what would happen and changes nothing off
TINYTITAN_DSH_NO_OPEN 1 never opens a browser window off

Diagnostics

These change no result; they make the engine report what it did. They are the first thing to reach for when something is slow or wrong.

Variable Effect
TINYTITAN_RUNNER_STATS Print the per-stage timing footer (prompt, embed, I/O, decode)
TINYTITAN_KERNEL_STATS Per-command-buffer GPU timings
TINYTITAN_KERNEL_SPLIT 1 gives every kernel in the decode chain its own command buffer
TINYTITAN_DECODE_IO_TRACE 1 reports the decode phase's own expert I/O
TINYTITAN_ROUTE_TRACE Names a file for the per-token routing trace (JSONL)
TINYTITAN_PREFETCH_TRACE Names a file for the predictive-prefetch trace
TINYTITAN_ACT_DUMP Names a directory to dump activations into
TINYTITAN_MTP_TRACE 1 traces each MTP draft/verify step on stderr
TINYTITAN_ANE_MEMORY_TRACE 1 traces the ANE runtime's arena footprint
TINYTITAN_LOGIT_TRACE 1 prints the top-2 logits per decode step
TINYTITAN_WATCHDOGS 1 enables the server's stall watchdogs

Other tools

  • tools/dsh_route.sh — print (or --write) the DeepSeek Harness route for the local server, generated from the installs you actually have: tools/dsh_route.sh, --write, --models qwen38 4, --reasoning off, --port, --context, --max-tokens, --provider, --from-server, --settings. Exits 2 when there is nothing it can describe.
  • tools/dsh_local.sh — TinyTitan's own isolated DeepSeek Harness, the optional browser window: ensure, web, smoke, port, status, paths. It keeps its own home, npm prefix, pnpm store and port, and never reads or writes a DeepSeek Harness you run yourself. See TinyTitan Plugin.
  • ttlanmanager — the LAN control plane for several Macs running DeepSeek Harness: top, list, prompt, prompt-all, workspace create, session archive, workspace delete, plus --peer HOST[:PORT], --key, --json, --timeout and --once/--width/--height/--from for rendering an inventory without a fleet. See LAN Manager.
  • TinyTitanRepack — the installer/converter behind the catalogue: --model <key> --output <dir> (with --resume, --overwrite), --input-snapshot <dir> --model-id <id>, --discard-partial --output <dir>, and --verify-install --input-gturbo <dir> to re-issue a receipt after a move.

Clone this wiki locally