Skip to content

InferKit v0.3.0

Choose a tag to compare

@belisoful belisoful released this 07 Sep 03:18
· 150 commits to main since this release

InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible and Anthropic
chat / vision / audio / embedding / image / video services (hosted, or a local runner such as Ollama
or LM Studio), audio transcription, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.

Two optional Swift companion packages build on the core without raising its floor:

  • InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — 60-plus real on-device models at measured
    reference parity: text-to-image and video generators (Stable Diffusion, Z-Image, SANA, LTX-Video,
    Wan), language models (Qwen3 dense and mixture-of-experts, Gemma 4, DeepSeek V4, and the hybrid
    Qwen3.5+ decoder) with a full generation runtime, text embeddings and reranking, vision-language
    models, speech recognition (Whisper, Parakeet) and speech synthesis / voice cloning (Kokoro,
    Chatterbox), a speech-restoration family, detection / segmentation / matting / depth, neural audio
    codecs, native GGUF and PyTorch checkpoint readers, runtime quantization, and on-device fine-tuning.
  • InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model, with
    streaming, tool calling, and structured output.

What's new in 0.3.0

  • The remote surface is complete — NFKRemoteBackend and NFKAnthropicBackend now stream, call
    tools, and return structured output, and take images, audio, documents, and sampled video in.
    Thirteen OpenAI-compatible providers plus Anthropic ship as presets, alongside backends for
    embeddings, reranking, moderation, transcription (with timestamps and translation), and speech /
    image / video generation, a model catalog, and native adapters for the local runners (Ollama,
    LM Studio).
  • On-device language models through MLX — NFKMLXLanguage runs Qwen3 (dense and Qwen3-MoE /
    Mixtral), Gemma 4 (the E-series, the 26B-A4B mixture, and the 12B unified), DeepSeek V4, and the
    hybrid Qwen3.5 / 3.6 / 3.8 decoder, each at measured reference parity — with a full generation
    runtime: key-value cache quantization and a bounded context window, prompt caching, speculative
    decoding, JSON- and choice-constrained decoding, and a native chat-template renderer.
  • Native checkpoint readers — a pure-Swift GGUF reader and a PyTorch .pth / .pt / .ckpt
    reader, so every model loads a raw released checkpoint with no Python toolchain.
  • Text embeddings, reranking, and vision-language models — Qwen3-Embedding and EmbeddingGemma,
    a ModernBERT cross-encoder reranker, and the SmolVLM and Qwen3-VL vision towers.
  • A wave of new on-device models at reference parity — Parakeet-TDT and Chatterbox voice cloning,
    Kokoro TTS, the speech-restoration / enhancement family (MP-SENet, GTCRN, SGMSE+, StoRM, MossFormer2
    SE, DeepFilterNet3, VoiceRestore, Resemble Enhance), RT-DETR and RF-DETR detection, BiRefNet matting,
    Depth Anything 3, the DAC and SNAC audio codecs, Silero VAD v6, SigLIP 2, TAESD, and the LTX-Video /
    Z-Image / SANA / Wan generation stacks.
  • Objective-C parity across the MLX surface, plus a DocC model gallery and index.

Full change log: CHANGELOG.md.

Upgrading from 0.2.x

Pre-1.0, a minor bump may change public API, and SwiftPM treats a 0.x minor as breaking — move your
dependency to 0.3.0 deliberately (from: "0.3.0" / ~> 0.3).

Install

Swift Package Manager (source, recommended):

.package(url: "https://github.com/belisoful/InferKit.git", from: "0.3.0")

CocoaPods (core): pod 'InferKit', '~> 0.3'

From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.

Binary assets — which to download

asset pick it when
InferKit.xcframework.zip (1.3 MB) you want the prebuilt core alone, no MLX
InferKitMLX.xcframework.zip (37 MB) most apps — static, core + MLX together, the best fit for an Objective-C consumer
InferKitMLXDynamic.xcframework.zip (24 MB) one shared framework across several targets or plug-ins

In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content
— static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.

As a SwiftPM binaryTarget:

.binaryTarget(name: "InferKitMLX",
              url: "https://github.com/belisoful/InferKit/releases/download/v0.3.0/InferKitMLX.xcframework.zip",
              checksum: "1c9d32bade85aeb2c0bf665ff8ed84eb9e05b7a5f535b4e010e8adbb8605e07f")

Checksums (swift package compute-checksum):

InferKit.xcframework.zip           b91b9070929e6ac9a0e580bb2f60267d6bb54a73ade19441f6c407d58201f6ea

InferKitMLX.xcframework.zip        1c9d32bade85aeb2c0bf665ff8ed84eb9e05b7a5f535b4e010e8adbb8605e07f

InferKitMLXDynamic.xcframework.zip b78c24c1a5b4902440e1a369620767ce494beae2b2e66630fb9fe8e0f4e8f049

The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.