Skip to content

FluidUse 0.3.0

Latest

Choose a tag to compare

@Alex-Wengg Alex-Wengg released this 25 Sep 18:52
· 6 commits to main since this release
a054c68

GLiNER2.5-Decide on Core ML

  • GLiNER2Variant.decide (128 tokens) and .decideLong (256 tokens): GLiNER2.5-Decide (Fastino, Apache-2.0) from FluidInference/gliner2-5-decide-coreml. Up to 4 questions × 32 labels per call; fp16 packages pinned by revision and checksum.
  • GLiNER2Manager.classifyConcurrently(text:heads:): several questions per call, and several calls in flight on one model.
  • SortAnythingDemo: 1,000 Wikipedia abstracts sorted into categories chosen at run time (Show and Turbo modes, SORT_AUTOPLAY for recordings).
  • SortDecisionsDemo: Fastino's Fast Decisions set, every question of a document answered in one call.

Numbers (Apple M5 Pro, 1,000 DBpedia abstracts)

Core ML (fp16) gliner2 PyTorch, Mac GPU
Time ~5.9 s 22.8 s
Peak memory ~1.0 GB 5.7 GB
Weights 0.92 GB 1.95 GB (fp32)
Accuracy 89.0% 89.0%

Identical predictions on all 1,000. Fast Decisions (public split): 62.9% average, identical to the original model on all 2,900 decisions.

Also since 0.2.0

  • GLiClass decision demos and 2048 model comparison (#6); GLiClass models load from Hugging Face (#7).
  • GLiNER 2.5 small / base / multilingual Core ML classifiers (#8).
  • Sub-1B decision-model runtimes (#9).
  • laya runtime and Tetris demo fixes (#5), package docs (#4).

Package

.package(url: "https://github.com/FluidInference/FluidUse.git", from: "0.3.0")

Pull requests: #4, #5, #6, #7, #8, #9, #13.