Skip to content

Releases: belisoful/InferKit

InferKit v0.3.1

Choose a tag to compare

@belisoful belisoful released this 16 Sep 03:49

InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible and Anthropic
chat / vision / audio / embedding / image / video services (hosted, or a local runner such as Ollama
or LM Studio), audio transcription, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.

Two optional Swift companion packages build on the core without raising its floor:

  • InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — a hundred-plus real on-device models at
    measured reference parity: text-to-image and video generators (Stable Diffusion 1.5 through 3.5,
    FLUX.1, Z-Image, SANA, LTX-Video, Wan), language models (Qwen3 dense and mixture-of-experts,
    gpt-oss, Gemma 3 / 3n / 4, DeepSeek V4, and the hybrid Qwen3.5+ decoder) with a full generation
    runtime, text embeddings and reranking, vision-language models, speech recognition (Whisper,
    Parakeet) and speech synthesis / voice cloning (Kokoro, Chatterbox), a speech-restoration family,
    detection / segmentation / matting / depth / pose, neural audio codecs, native GGUF and PyTorch
    checkpoint readers, runtime quantization, and on-device fine-tuning.
  • InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model, with
    streaming, tool calling, and structured output.

What's new in 0.3.1

  • On-device fine-tuning is reachable from an app. NFKMLXWeights, NFKMLXQuantization and
    NFKMLXError are public, so the whole path — build a network, train, save, reload through the
    model's own factory — compiles for a consumer who links the package as an ordinary dependency. It
    previously compiled only inside the package's own tests.
  • A frozen backbone stays frozen. A training run now returns every fully frozen subtree to
    evaluation mode, so a head-only fine-tune over a pretrained convolutional backbone no longer
    normalizes with its own batches and overwrites the released statistics in the saved checkpoint.
    A run with nothing trainable throws instead of reporting a loss curve for an update that changed
    nothing.
  • Stable Diffusion 3 / 3.5 and FLUX.1, with both ControlNets, at reference parity — the MMDiT
    dual-stream transformer, FLUX's double- and single-stream blocks, and the partial-copy ControlNets
    whose per-block residuals the base transformers inject.
  • Gemma 3 and Gemma 3n at released-weight parity: Gemma 3 at 270M, 1B and 4B including its vision
    tower, and the tri-modal Gemma 3n (E2B and E4B) with its AltUp / LAuReL / per-layer-embedding
    decoder, USM audio encoder, and MobileNetV5 vision tower.
  • Every YOLO generation after v8 — YOLOv9 through YOLO26, 26 checkpoints read through one graph
    interpreter — plus ViTPose, DDColor, RT-DETRv2, and the successors of the older shipped image
    models: IS-Net, HAT, AdaIN, Zero-DCE++, the Real-ESRGAN compact releases, and MetaCLIP weights.
  • Six more speech models: MetricGAN+, CMGAN, FRCRN, MossFormer2 SR, NU-Wave 2 and Apollo, each at
    reference parity on its first numeric run.
  • Schema-constrained decoding, on device and engine-agnostic: a JSON Schema compiles to a
    byte-level grammar that masks the sampler, and the core carries JSON and fixed-choice grammars so
    the Core ML language backend constrains its own output too.
  • gpt-oss (MXFP4) and Qwen2-MoE, a per-channel key-value cache that makes 8-bit quantization
    near-lossless, Depth Anything 3's ray and camera branches, SmolVLM2's other sizes, and full released
    size coverage across the shipped families.
  • Local-runner discovery in the core — NFKRemoteProvider probes the local presets and answers
    with the ones that reply, so calling code no longer names the app the user happens to be running.
  • Builds on Xcode 27. InferKitMLX compiles in Release again, and the MLX tests run under plain
    swift test.

Full change log: CHANGELOG.md.

Upgrading from 0.3.0

A patch release: no public API was removed, so a dependency on from: "0.3.0" picks this up. Three
things change behavior rather than signatures.

  • The quantized key-value cache groups keys per channel by default. A run at the same bit width and
    group size produces different logits, and 8-bit now tracks full precision closely.
    NFKMLXKeyValueCache.Quantization.keyAxis selects the previous per-position layout.
  • A fine-tune over a frozen BatchNorm backbone produces different weights than before, with the
    backbone's released statistics intact.
  • NFKMLXError is public and gains cases as the package grows, so a switch over it wants
    @unknown default.

Building InferKitMLX from source on Xcode 27 needs Xcode's on-demand Metal compiler, which
SwiftPM invokes during the build: xcodebuild -downloadComponent MetalToolchain.

Install

Swift Package Manager (source, recommended):

.package(url: "https://github.com/belisoful/InferKit.git", from: "0.3.1")

CocoaPods (core): pod 'InferKit', '~> 0.3'

From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.

Binary assets — which to download

asset pick it when
InferKit.xcframework.zip (1.3 MB) you want the prebuilt core alone, no MLX
InferKitMLX.xcframework.zip (38 MB) most apps — static, core + MLX together, the best fit for an Objective-C consumer
InferKitMLXDynamic.xcframework.zip (25 MB) one shared framework across several targets or plug-ins

In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content
— static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.

As a SwiftPM binaryTarget:

.binaryTarget(name: "InferKitMLX",
              url: "https://github.com/belisoful/InferKit/releases/download/v0.3.1/InferKitMLX.xcframework.zip",
              checksum: "1e104f6a0d99a5e5f3d28a9b19fdc1927cc9a26ee00ce8686ea868c156640681")

Checksums (swift package compute-checksum):

InferKit.xcframework.zip           b384a8da0632a113c4f3355a81e159cc9cf8d3a960eb5e8a9132a53fc8113ec0

InferKitMLX.xcframework.zip        1e104f6a0d99a5e5f3d28a9b19fdc1927cc9a26ee00ce8686ea868c156640681

InferKitMLXDynamic.xcframework.zip e15bc05cbc98daa13bbdf32af5a9889a0048daecac25698db840f37e16d27efd

The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.

InferKit v0.3.0

Choose a tag to compare

@belisoful belisoful released this 07 Sep 03:18

InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible and Anthropic
chat / vision / audio / embedding / image / video services (hosted, or a local runner such as Ollama
or LM Studio), audio transcription, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.

Two optional Swift companion packages build on the core without raising its floor:

  • InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — 60-plus real on-device models at measured
    reference parity: text-to-image and video generators (Stable Diffusion, Z-Image, SANA, LTX-Video,
    Wan), language models (Qwen3 dense and mixture-of-experts, Gemma 4, DeepSeek V4, and the hybrid
    Qwen3.5+ decoder) with a full generation runtime, text embeddings and reranking, vision-language
    models, speech recognition (Whisper, Parakeet) and speech synthesis / voice cloning (Kokoro,
    Chatterbox), a speech-restoration family, detection / segmentation / matting / depth, neural audio
    codecs, native GGUF and PyTorch checkpoint readers, runtime quantization, and on-device fine-tuning.
  • InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model, with
    streaming, tool calling, and structured output.

What's new in 0.3.0

  • The remote surface is complete — NFKRemoteBackend and NFKAnthropicBackend now stream, call
    tools, and return structured output, and take images, audio, documents, and sampled video in.
    Thirteen OpenAI-compatible providers plus Anthropic ship as presets, alongside backends for
    embeddings, reranking, moderation, transcription (with timestamps and translation), and speech /
    image / video generation, a model catalog, and native adapters for the local runners (Ollama,
    LM Studio).
  • On-device language models through MLX — NFKMLXLanguage runs Qwen3 (dense and Qwen3-MoE /
    Mixtral), Gemma 4 (the E-series, the 26B-A4B mixture, and the 12B unified), DeepSeek V4, and the
    hybrid Qwen3.5 / 3.6 / 3.8 decoder, each at measured reference parity — with a full generation
    runtime: key-value cache quantization and a bounded context window, prompt caching, speculative
    decoding, JSON- and choice-constrained decoding, and a native chat-template renderer.
  • Native checkpoint readers — a pure-Swift GGUF reader and a PyTorch .pth / .pt / .ckpt
    reader, so every model loads a raw released checkpoint with no Python toolchain.
  • Text embeddings, reranking, and vision-language models — Qwen3-Embedding and EmbeddingGemma,
    a ModernBERT cross-encoder reranker, and the SmolVLM and Qwen3-VL vision towers.
  • A wave of new on-device models at reference parity — Parakeet-TDT and Chatterbox voice cloning,
    Kokoro TTS, the speech-restoration / enhancement family (MP-SENet, GTCRN, SGMSE+, StoRM, MossFormer2
    SE, DeepFilterNet3, VoiceRestore, Resemble Enhance), RT-DETR and RF-DETR detection, BiRefNet matting,
    Depth Anything 3, the DAC and SNAC audio codecs, Silero VAD v6, SigLIP 2, TAESD, and the LTX-Video /
    Z-Image / SANA / Wan generation stacks.
  • Objective-C parity across the MLX surface, plus a DocC model gallery and index.

Full change log: CHANGELOG.md.

Upgrading from 0.2.x

Pre-1.0, a minor bump may change public API, and SwiftPM treats a 0.x minor as breaking — move your
dependency to 0.3.0 deliberately (from: "0.3.0" / ~> 0.3).

Install

Swift Package Manager (source, recommended):

.package(url: "https://github.com/belisoful/InferKit.git", from: "0.3.0")

CocoaPods (core): pod 'InferKit', '~> 0.3'

From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.

Binary assets — which to download

asset pick it when
InferKit.xcframework.zip (1.3 MB) you want the prebuilt core alone, no MLX
InferKitMLX.xcframework.zip (37 MB) most apps — static, core + MLX together, the best fit for an Objective-C consumer
InferKitMLXDynamic.xcframework.zip (24 MB) one shared framework across several targets or plug-ins

In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content
— static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.

As a SwiftPM binaryTarget:

.binaryTarget(name: "InferKitMLX",
              url: "https://github.com/belisoful/InferKit/releases/download/v0.3.0/InferKitMLX.xcframework.zip",
              checksum: "1c9d32bade85aeb2c0bf665ff8ed84eb9e05b7a5f535b4e010e8adbb8605e07f")

Checksums (swift package compute-checksum):

InferKit.xcframework.zip           b91b9070929e6ac9a0e580bb2f60267d6bb54a73ade19441f6c407d58201f6ea

InferKitMLX.xcframework.zip        1c9d32bade85aeb2c0bf665ff8ed84eb9e05b7a5f535b4e010e8adbb8605e07f

InferKitMLXDynamic.xcframework.zip b78c24c1a5b4902440e1a369620767ce494beae2b2e66630fb9fe8e0f4e8f049

The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.

InferKit v0.2.0

Choose a tag to compare

@belisoful belisoful released this 29 Aug 11:32

InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible chat and
transcription services, Anthropic, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.

Two optional Swift companion packages build on the core without raising its floor:

  • InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — 35+ real on-device models at measured
    reference parity: Stable Diffusion, Whisper (with timestamps), SAM 2, YOLOv8, Demucs, MiniMax
    Music 3 (text + lyrics → music), a complete trained text-to-speech voice, language models (Qwen3,
    Qwen3.5, Gemma 4), runtime quantization, and on-device fine-tuning (trainers, LoRA).
  • InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model.

What's new in 0.2.0

  • MiniMax Music 3 — NFKMLXMusic3 / NFKMLXMusicBackend: a music description plus lyrics
    generate stereo 44.1 kHz music, its five networks each at measured reference parity.
  • Runtime MLX quantization — NFKMLXQuantization packs a model to affine 4-/8-bit;
    NFKMLXMusic3.quantizeRelease writes a quantized release (4-bit language model including its untied
    input embedding, 8-bit DiT), taking the stack from 27 GB to 7.7 GiB and letting it stay resident.
  • Qwen tokenization fix — the Qwen text path now uses qwen2 pre-tokenization; the byte-level
    BPE GPT-2 default encoded the same prompt to different, valid-looking token ids.
  • NFKInferKit.version — the core reports its version string.
  • Docs — DocC pages for the music classes.

Full change log: CHANGELOG.md.

Upgrading from 0.1.x

Pre-1.0, a minor bump may change public API, and SwiftPM treats a 0.x minor as breaking — move your
dependency to 0.2.0 deliberately (from: "0.2.0" / ~> 0.2).

Install

Swift Package Manager (source, recommended):

.package(url: "https://github.com/belisoful/InferKit.git", from: "0.2.0")

CocoaPods (core): pod 'InferKit', '~> 0.2'

From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.

Binary assets — which to download

asset pick it when
InferKit.xcframework.zip (0.8 MB) you want the prebuilt core alone, no MLX
InferKitMLX.xcframework.zip (29 MB) most apps — static, core + MLX together, the best fit for an Objective-C consumer
InferKitMLXDynamic.xcframework.zip (19 MB) one shared framework across several targets or plug-ins

In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content
— static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.

As a SwiftPM binaryTarget:

.binaryTarget(name: "InferKitMLX",
              url: "https://github.com/belisoful/InferKit/releases/download/v0.2.0/InferKitMLX.xcframework.zip",
              checksum: "fd08f47fcc37af1ee46a0216c0b9f7111b9e3232c079f99f9961d7aab062c986")

Checksums (swift package compute-checksum):

InferKit.xcframework.zip           8c8632ab0df7d8a4ef66720025f17c1088b21b9b4075a6869a62b30ff4188b5c

InferKitMLX.xcframework.zip        fd08f47fcc37af1ee46a0216c0b9f7111b9e3232c079f99f9961d7aab062c986

InferKitMLXDynamic.xcframework.zip 9ed0f6ffa6dc3b86670b97b77cbd3188a52925df328acc03487dacae51fb9cdc

The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.

InferKit v0.1.0

Choose a tag to compare

@belisoful belisoful released this 23 Aug 03:22

InferKit is a cross-platform inference toolkit for Objective-C (class prefix NFK): a swappable
backend protocol, immutable request/result value types, an async job handle, and shipped backends for
Core ML (including an on-device causal-language-model runner), OpenAI-compatible chat and
transcription services, Anthropic, and runtime backend discovery. It has no host-framework
dependency, so any Metal/Apple app can use it. Core floor: macOS 11 / iOS 14 / tvOS 14. MIT license.

Two optional Swift companion packages build on the core without raising its floor:

  • InferKitMLX (Apple Silicon, macOS 14 / iOS 17) — 30+ real on-device models at measured
    reference parity: Stable Diffusion, Whisper (with timestamps), SAM 2, YOLOv8, Demucs, a complete
    trained text-to-speech voice, language models (Qwen3, Qwen3.5, Gemma 4), and on-device
    fine-tuning (trainers, LoRA).
  • InferKitFoundationModels (macOS 26 / iOS 26) — Apple's on-device system language model.

Install

Swift Package Manager (source, recommended):

.package(url: "https://github.com/belisoful/InferKit.git", from: "0.1.0")

CocoaPods (core): pod 'InferKit', '~> 0.1'

From source: swift build && swift test at the root; cd InferKitMLX && swift build for the
MLX companion.

Binary assets — which to download

asset pick it when
InferKit.xcframework.zip (0.7 MB) you want the prebuilt core alone, no MLX
InferKitMLX.xcframework.zip (28 MB) most apps — static, core + MLX together, the best fit for an Objective-C consumer
InferKitMLXDynamic.xcframework.zip (19 MB) one shared framework across several targets or plug-ins

In Xcode: unzip, then drag the .xcframework into your target's Frameworks, Libraries, and
Embedded Content
— static: Do Not Embed; dynamic: Embed & Sign. The static MLX slices carry the
mlx-swift_Cmlx.bundle Metal shaders: add that bundle to Copy Bundle Resources, or MLX throws
Failed to load the default metallib at the first inference.

As a SwiftPM binaryTarget:

.binaryTarget(name: "InferKitMLX",
              url: "https://github.com/belisoful/InferKit/releases/download/v0.1.0/InferKitMLX.xcframework.zip",
              checksum: "cd7d0fc1acef39406d7c0c54bbe9c3eb8ea5f8379979f69d8201e9deb6eeb9c9")

Checksums (swift package compute-checksum):

InferKit.xcframework.zip           acc6d333d83dfdfb3190fe9676bcadd376a33ec4017c77aeb9f6be141a5b45c9

InferKitMLX.xcframework.zip        cd7d0fc1acef39406d7c0c54bbe9c3eb8ea5f8379979f69d8201e9deb6eeb9c9

InferKitMLXDynamic.xcframework.zip 69a75248c07f20ea449801cdb9fe24590d2aab0dc70ee1857e383636318b15fd

The recipe for every consumer shape — plug-in bundles, the dynamic variant's CoreHeaders
modulemap for Objective-C, linking from your own static library — is in
Docs/installation.md.
Examples: Docs/examples.md.
Full change log: CHANGELOG.md.