Skip to content

Library and Engine

André Borchert edited this page Oct 3, 2026 · 12 revisions
TinyTitan

Library and engine

TinyTitan ships two products on one branch. They are deliberately separate, and they are separate the same way the two DeepSeek Harness bundles are separate under plugins/: one repository, one release tag, one CI, and products with different surfaces. docs/plan-embedded-library.md is the design record; docs/repository-layout.md maps the targets.

Product What you do with it
TinyTitanLib — the library Add it as a Swift package dependency and embed the engine in your own program.
The engine — TinyTitanCLI, TinyTitanServer, the installer and the reporter Install it and run models from the terminal.

The engine is built on the library, not beside it: the server serves generation through the library's own session, so there is one implementation of loading, prompt rendering and sampling.

The library

TinyTitanLib is a Swift package product. Add the package and depend on the library:

.package(url: "https://github.com/Pummelchen/TinyTitan", from: "5.15.0")
import TinyTitanLib

let device = MTLCreateSystemDefaultDevice()!
let engine = try await Engine(directory: installURL, device: device)
let session = await engine.session()

let summary = try await session.respond(
    to: [ChatMessage(role: .user, content: "The capital of France is")],
    options: GenerationOptions(maxTokens: 32, temperature: 0)
) { event in
    if case .token(let text) = event { print(text, terminator: "") }
}
print(summary.stopReason, summary.completionTokens)

That is the whole shape: an Engine holds one install and the resident weights, a Session holds one conversation, and respond streams tokens through a callback. Four types do the work — EngineConfiguration, ChatMessage, GenerationOptions and the GenerationEvent/GenerationSummary pair — and TinyTitanError names the failures a caller can act on.

There is no subprocess and no HTTP in that path. It is the same kernels, format reader, KV cache and sampler the installed engine uses, in your process.

The library logs to stderr, never to stdout, so it cannot corrupt what your program prints. If you own your output, take the lines as well: EngineConfiguration(logSink:) routes every diagnostic into a closure of yours, and { _ in } silences the library completely — which is what the CLI's --quiet does. The destination is process-wide, because the orchestrator logs statically from deep inside the engine; the most recently created Engine sets it.

Try it without writing anything. examples/embedded in the source repository is a consumer package with its own Package.swift that builds against the released tag; tools/embedded-dependency-check.sh builds it and, on a machine with an install, streams tokens from it, so the dependency cannot rot silently.

Note

The library needs Apple Silicon, macOS 26 or later, and Xcode 27 with Swift 6.4 — the only supported toolchain; nothing here is tested on another Swift — plus a .ssdai install made by the installer or tools/install_models.sh. It never downloads or ships weights, and model licensing stays yours.

The binary form

Every release carries the library as binaries too, in a second archive beside the engine's: tinytitan-lib-<version>-macos-arm64.tar.gz. It holds libTinyTitanLib.a, libTinyTitanLib.dylib, the Swift module and every module it imports, the generated module maps, the resource bundle, the licences, a README-library.txt with these commands, and demo/ — two terminal apps built from one source, one per link form, so there is a working example to read and run:

cd demo && ./build-static.sh && ./tinytitan-demo-static --model <install>
./make-flags.sh          # names the module maps of this directory
swiftc -I . @swiftc-flags.txt mytool.swift libTinyTitanLib.a -o mytool   # static
swiftc -I . @swiftc-flags.txt -L . -lTinyTitanLib mytool.swift \
  -Xlinker -rpath -Xlinker "$PWD" -o mytool                              # dynamic

Two things that are easy to get wrong, and that the archive handles:

  • The static archive is merged. libTinyTitanLib.a contains the library and every dependency it needs, so it is the one file you link.
  • The Metal shaders travel as sources, in TinyTitan_TinyTitan.bundle, and are compiled when the library loads. Keep the bundles next to your executable, or the runtime cannot find its kernels.

Building the package as a SwiftPM dependency remains the supported route and needs none of those flags; this archive is for a consumer that cannot or will not build from source. Its module is built for Xcode 27 / Swift 6.4 and is not stable across other toolchains — which costs nothing here, because no other toolchain is supported.

The engine

The engine is what the Getting Started guide installs: the tinytitan command, the CLI, the repacker, the benchmark driver, and the loopback server described in Local Server and API. It is the product for running a model on this Mac; the library is the product for building your own front end on it. Both front ends generate through the library — the CLI's output is byte-identical to what it produced before the move, checked against saved baselines — so there is one generation path to maintain, not two. Neither product downloads weights, and neither is a network service — the server binds 127.0.0.1 only.

What the library does not do yet

The facade is new and deliberately small. These are known and recorded in docs/plan-embedded-library.md rather than hidden:

  • The device you pass is the device it runs on. Engine(directory:device:) builds the Metal context, its queues and its pipelines on that device — so a machine with more than one GPU, or a caller holding a device of its own, is not quietly redirected to the system default.
  • The streaming mode is not a knob, because it is not a choice — the runtime has one mode, and the slot count it takes is expertCacheSlots. What is a knob is EngineConfiguration.integrityPolicy: trust the installer receipt, re-hash every file, or let the loader decide as it always has.
  • The facade's events carry everything the orchestrator produces — promptProcessed, token, reasoning, toolCall, finished — and the summary repeats reasoning and toolCalls.
  • Tools are part of the loop, both directions. Pass tools: to respond with each schema as JSON text, read the model's toolCall events, then replay the assistant turn that asked (with its toolCalls) plus a tool message carrying the result and its toolCallID. A tool result that names no open call is refused, which is why the assistant turn has to go back in too.
  • A local caller is not held to OpenAI's limits. The wire's caps — four stop strings, a thousand messages — apply to what arrives over HTTP; a caller inside the process passes its own, and the structural rules (a message needs content, an unknown role is refused) apply to everyone.
  • One generation per session, and a second one waits. That is the contract, not a limitation: the engine's slot pool serialises them, the outputs both come back intact, and tests/TinyTitanLib/LibraryContractTests.swift pins it along with two engines on one device, cancellation and unload().
  • Requests are validated the way the server validates them. A conversation is checked against the OpenAI wire rules (message count, stop-string count, a tool role needing a tool-call id), which is stricter than a local caller may expect.

Anything not listed in the facade's public types is not a supported surface: the rest of TinyTitanLib is package, which is what makes the promise above possible.

Versioning

Both products ship from one tag, and the library's version is that tag — there is no second version to keep in step. Until the release that introduces it says so, main is the reference for the facade and the plan is the record of what is settled.

Clone this wiki locally