Skip to content

InferKit v0.2.0

Choose a tag to compare

@belisoful belisoful released this 29 Aug 11:32
· 111 commits to main since this release

InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible chat and
transcription services, Anthropic, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.

Two optional Swift companion packages build on the core without raising its floor:

  • InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — 35+ real on-device models at measured
    reference parity: Stable Diffusion, Whisper (with timestamps), SAM 2, YOLOv8, Demucs, MiniMax
    Music 3 (text + lyrics → music), a complete trained text-to-speech voice, language models (Qwen3,
    Qwen3.5, Gemma 4), runtime quantization, and on-device fine-tuning (trainers, LoRA).
  • InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model.

What's new in 0.2.0

  • MiniMax Music 3 — NFKMLXMusic3 / NFKMLXMusicBackend: a music description plus lyrics
    generate stereo 44.1 kHz music, its five networks each at measured reference parity.
  • Runtime MLX quantization — NFKMLXQuantization packs a model to affine 4-/8-bit;
    NFKMLXMusic3.quantizeRelease writes a quantized release (4-bit language model including its untied
    input embedding, 8-bit DiT), taking the stack from 27 GB to 7.7 GiB and letting it stay resident.
  • Qwen tokenization fix — the Qwen text path now uses qwen2 pre-tokenization; the byte-level
    BPE GPT-2 default encoded the same prompt to different, valid-looking token ids.
  • NFKInferKit.version — the core reports its version string.
  • Docs — DocC pages for the music classes.

Full change log: CHANGELOG.md.

Upgrading from 0.1.x

Pre-1.0, a minor bump may change public API, and SwiftPM treats a 0.x minor as breaking — move your
dependency to 0.2.0 deliberately (from: "0.2.0" / ~> 0.2).

Install

Swift Package Manager (source, recommended):

.package(url: "https://github.com/belisoful/InferKit.git", from: "0.2.0")

CocoaPods (core): pod 'InferKit', '~> 0.2'

From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.

Binary assets — which to download

asset pick it when
InferKit.xcframework.zip (0.8 MB) you want the prebuilt core alone, no MLX
InferKitMLX.xcframework.zip (29 MB) most apps — static, core + MLX together, the best fit for an Objective-C consumer
InferKitMLXDynamic.xcframework.zip (19 MB) one shared framework across several targets or plug-ins

In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content
— static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.

As a SwiftPM binaryTarget:

.binaryTarget(name: "InferKitMLX",
              url: "https://github.com/belisoful/InferKit/releases/download/v0.2.0/InferKitMLX.xcframework.zip",
              checksum: "fd08f47fcc37af1ee46a0216c0b9f7111b9e3232c079f99f9961d7aab062c986")

Checksums (swift package compute-checksum):

InferKit.xcframework.zip           8c8632ab0df7d8a4ef66720025f17c1088b21b9b4075a6869a62b30ff4188b5c

InferKitMLX.xcframework.zip        fd08f47fcc37af1ee46a0216c0b9f7111b9e3232c079f99f9961d7aab062c986

InferKitMLXDynamic.xcframework.zip 9ed0f6ffa6dc3b86670b97b77cbd3188a52925df328acc03487dacae51fb9cdc

The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.