-
Notifications
You must be signed in to change notification settings - Fork 2
Library and Engine
TinyTitan ships two products on one branch. They are deliberately separate,
and they are separate the same way the two DeepSeek Harness bundles are separate
under plugins/: one repository, one release tag, one CI, and products with
different surfaces. docs/plan-embedded-library.md is the design record;
docs/repository-layout.md maps the targets.
| Product | What you do with it |
|---|---|
TinyTitanLib — the library |
Add it as a Swift package dependency and embed the engine in your own program. |
The engine — TinyTitanCLI, TinyTitanServer, the installer and the reporter |
Install it and run models from the terminal. |
The engine is built on the library, not beside it: the server serves generation through the library's own session, so there is one implementation of loading, prompt rendering and sampling.
TinyTitanLib is a Swift package product. Add the package and depend on the
library:
.package(url: "https://github.com/Pummelchen/TinyTitan", from: "5.15.0")import TinyTitanLib
let device = MTLCreateSystemDefaultDevice()!
let engine = try await Engine(directory: installURL, device: device)
let session = await engine.session()
let summary = try await session.respond(
to: [ChatMessage(role: .user, content: "The capital of France is")],
options: GenerationOptions(maxTokens: 32, temperature: 0)
) { event in
if case .token(let text) = event { print(text, terminator: "") }
}
print(summary.stopReason, summary.completionTokens)That is the whole shape: an Engine holds one install and the resident weights,
a Session holds one conversation, and respond streams tokens through a
callback. Four types do the work — EngineConfiguration, ChatMessage,
GenerationOptions and the GenerationEvent/GenerationSummary pair — and
TinyTitanError names the failures a caller can act on.
There is no subprocess and no HTTP in that path. It is the same kernels, format reader, KV cache and sampler the installed engine uses, in your process.
The library logs to stderr, never to stdout, so it cannot corrupt what your
program prints. If you own your output, take the lines as well:
EngineConfiguration(logSink:) routes every diagnostic into a closure of yours,
and { _ in } silences the library completely — which is what the CLI's
--quiet does. The destination is process-wide, because the orchestrator logs
statically from deep inside the engine; the most recently created Engine sets
it.
Try it without writing anything. examples/embedded in the source repository
is a consumer package with its own Package.swift that builds against the
released tag; tools/embedded-dependency-check.sh builds it and, on a machine
with an install, streams tokens from it, so the dependency cannot rot silently.
Note
The library needs Apple Silicon, macOS 26 or later, and Xcode 27 with Swift
6.4 — the only supported toolchain; nothing here is tested on another Swift
— plus a .ssdai install made by the installer or tools/install_models.sh.
It never downloads or ships weights, and model licensing stays yours.
Every release carries the library as binaries too, in a second archive beside
the engine's: tinytitan-lib-<version>-macos-arm64.tar.gz. It holds
libTinyTitanLib.a, libTinyTitanLib.dylib, the Swift module and every
module it imports, the generated module maps, the resource bundle, the
licences, a README-library.txt with these commands, and demo/ — two terminal
apps built from one source, one per link form, so there is a working example to
read and run:
cd demo && ./build-static.sh && ./tinytitan-demo-static --model <install>./make-flags.sh # names the module maps of this directory
swiftc -I . @swiftc-flags.txt mytool.swift libTinyTitanLib.a -o mytool # static
swiftc -I . @swiftc-flags.txt -L . -lTinyTitanLib mytool.swift \
-Xlinker -rpath -Xlinker "$PWD" -o mytool # dynamicTwo things that are easy to get wrong, and that the archive handles:
-
The static archive is merged.
libTinyTitanLib.acontains the library and every dependency it needs, so it is the one file you link. -
The Metal shaders travel as sources, in
TinyTitan_TinyTitan.bundle, and are compiled when the library loads. Keep the bundles next to your executable, or the runtime cannot find its kernels.
Building the package as a SwiftPM dependency remains the supported route and needs none of those flags; this archive is for a consumer that cannot or will not build from source. Its module is built for Xcode 27 / Swift 6.4 and is not stable across other toolchains — which costs nothing here, because no other toolchain is supported.
The engine is what the Getting Started guide installs: the
tinytitan command, the CLI, the repacker, the benchmark driver, and the
loopback server described in Local Server and API. It
is the product for running a model on this Mac; the library is the product for
building your own front end on it. Both front ends generate through the
library — the CLI's output is byte-identical to what it produced before the move,
checked against saved baselines — so there is one generation path to maintain,
not two. Neither product downloads weights, and neither is a network service —
the server binds 127.0.0.1 only.
The facade is new and deliberately small. These are known and recorded in
docs/plan-embedded-library.md rather than hidden:
-
The device you pass is the device it runs on.
Engine(directory:device:)builds the Metal context, its queues and its pipelines on that device — so a machine with more than one GPU, or a caller holding a device of its own, is not quietly redirected to the system default. -
The streaming mode is not a knob, because it is not a choice — the runtime
has one mode, and the slot count it takes is
expertCacheSlots. What is a knob isEngineConfiguration.integrityPolicy: trust the installer receipt, re-hash every file, or let the loader decide as it always has. -
The facade's events carry everything the orchestrator produces —
promptProcessed,token,reasoning,toolCall,finished— and the summary repeatsreasoningandtoolCalls. -
Tools are part of the loop, both directions. Pass
tools:torespondwith each schema as JSON text, read the model'stoolCallevents, then replay the assistant turn that asked (with itstoolCalls) plus atoolmessage carrying the result and itstoolCallID. A tool result that names no open call is refused, which is why the assistant turn has to go back in too. - A local caller is not held to OpenAI's limits. The wire's caps — four stop strings, a thousand messages — apply to what arrives over HTTP; a caller inside the process passes its own, and the structural rules (a message needs content, an unknown role is refused) apply to everyone.
-
One generation per session, and a second one waits. That is the contract,
not a limitation: the engine's slot pool serialises them, the outputs both come
back intact, and
tests/TinyTitanLib/LibraryContractTests.swiftpins it along with two engines on one device, cancellation andunload(). -
Requests are validated the way the server validates them. A conversation
is checked against the OpenAI wire rules (message count, stop-string count, a
toolrole needing a tool-call id), which is stricter than a local caller may expect.
Anything not listed in the facade's public types is not a supported surface: the
rest of TinyTitanLib is package, which is what makes the promise above
possible.
Both products ship from one tag, and the library's version is that tag —
there is no second version to keep in step. Until the release that introduces it
says so, main is the reference for the facade and the plan is the record of
what is settled.
Start
Use TinyTitan
DeepSeek Harness
Reference
Project