InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible and Anthropic
chat / vision / audio / embedding / image / video services (hosted, or a local runner such as Ollama
or LM Studio), audio transcription, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.
Two optional Swift companion packages build on the core without raising its floor:
- InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — a hundred-plus real on-device models at
measured reference parity: text-to-image and video generators (Stable Diffusion 1.5 through 3.5,
FLUX.1, Z-Image, SANA, LTX-Video, Wan), language models (Qwen3 dense and mixture-of-experts,
gpt-oss, Gemma 3 / 3n / 4, DeepSeek V4, and the hybrid Qwen3.5+ decoder) with a full generation
runtime, text embeddings and reranking, vision-language models, speech recognition (Whisper,
Parakeet) and speech synthesis / voice cloning (Kokoro, Chatterbox), a speech-restoration family,
detection / segmentation / matting / depth / pose, neural audio codecs, native GGUF and PyTorch
checkpoint readers, runtime quantization, and on-device fine-tuning. - InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model, with
streaming, tool calling, and structured output.
What's new in 0.3.1
- On-device fine-tuning is reachable from an app.
NFKMLXWeights,NFKMLXQuantizationand
NFKMLXErrorare public, so the whole path — build a network, train, save, reload through the
model's own factory — compiles for a consumer who links the package as an ordinary dependency. It
previously compiled only inside the package's own tests. - A frozen backbone stays frozen. A training run now returns every fully frozen subtree to
evaluation mode, so a head-only fine-tune over a pretrained convolutional backbone no longer
normalizes with its own batches and overwrites the released statistics in the saved checkpoint.
A run with nothing trainable throws instead of reporting a loss curve for an update that changed
nothing. - Stable Diffusion 3 / 3.5 and FLUX.1, with both ControlNets, at reference parity — the MMDiT
dual-stream transformer, FLUX's double- and single-stream blocks, and the partial-copy ControlNets
whose per-block residuals the base transformers inject. - Gemma 3 and Gemma 3n at released-weight parity: Gemma 3 at 270M, 1B and 4B including its vision
tower, and the tri-modal Gemma 3n (E2B and E4B) with its AltUp / LAuReL / per-layer-embedding
decoder, USM audio encoder, and MobileNetV5 vision tower. - Every YOLO generation after v8 — YOLOv9 through YOLO26, 26 checkpoints read through one graph
interpreter — plus ViTPose, DDColor, RT-DETRv2, and the successors of the older shipped image
models: IS-Net, HAT, AdaIN, Zero-DCE++, the Real-ESRGAN compact releases, and MetaCLIP weights. - Six more speech models: MetricGAN+, CMGAN, FRCRN, MossFormer2 SR, NU-Wave 2 and Apollo, each at
reference parity on its first numeric run. - Schema-constrained decoding, on device and engine-agnostic: a JSON Schema compiles to a
byte-level grammar that masks the sampler, and the core carries JSON and fixed-choice grammars so
the Core ML language backend constrains its own output too. - gpt-oss (MXFP4) and Qwen2-MoE, a per-channel key-value cache that makes 8-bit quantization
near-lossless, Depth Anything 3's ray and camera branches, SmolVLM2's other sizes, and full released
size coverage across the shipped families. - Local-runner discovery in the core —
NFKRemoteProviderprobes the local presets and answers
with the ones that reply, so calling code no longer names the app the user happens to be running. - Builds on Xcode 27. InferKitMLX compiles in Release again, and the MLX tests run under plain
swift test.
Full change log: CHANGELOG.md.
Upgrading from 0.3.0
A patch release: no public API was removed, so a dependency on from: "0.3.0" picks this up. Three
things change behavior rather than signatures.
- The quantized key-value cache groups keys per channel by default. A run at the same bit width and
group size produces different logits, and 8-bit now tracks full precision closely.
NFKMLXKeyValueCache.Quantization.keyAxisselects the previous per-position layout. - A fine-tune over a frozen
BatchNormbackbone produces different weights than before, with the
backbone's released statistics intact. NFKMLXErroris public and gains cases as the package grows, so aswitchover it wants
@unknown default.
Building InferKitMLX from source on Xcode 27 needs Xcode's on-demand Metal compiler, which
SwiftPM invokes during the build: xcodebuild -downloadComponent MetalToolchain.
Install
Swift Package Manager (source, recommended):
.package(url: "https://github.com/belisoful/InferKit.git", from: "0.3.1")CocoaPods (core): pod 'InferKit', '~> 0.3'
From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.
Binary assets — which to download
| asset | pick it when |
|---|---|
InferKit.xcframework.zip (1.3 MB) |
you want the prebuilt core alone, no MLX |
InferKitMLX.xcframework.zip (38 MB) |
most apps — static, core + MLX together, the best fit for an Objective-C consumer |
InferKitMLXDynamic.xcframework.zip (25 MB) |
one shared framework across several targets or plug-ins |
In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content — static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.
As a SwiftPM binaryTarget:
.binaryTarget(name: "InferKitMLX",
url: "https://github.com/belisoful/InferKit/releases/download/v0.3.1/InferKitMLX.xcframework.zip",
checksum: "1e104f6a0d99a5e5f3d28a9b19fdc1927cc9a26ee00ce8686ea868c156640681")Checksums (swift package compute-checksum):
InferKit.xcframework.zip b384a8da0632a113c4f3355a81e159cc9cf8d3a960eb5e8a9132a53fc8113ec0
InferKitMLX.xcframework.zip 1e104f6a0d99a5e5f3d28a9b19fdc1927cc9a26ee00ce8686ea868c156640681
InferKitMLXDynamic.xcframework.zip e15bc05cbc98daa13bbdf32af5a9889a0048daecac25698db840f37e16d27efd
The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.