GLiNER2.5-Decide on Core ML
GLiNER2Variant.decide(128 tokens) and.decideLong(256 tokens): GLiNER2.5-Decide (Fastino, Apache-2.0) from FluidInference/gliner2-5-decide-coreml. Up to 4 questions × 32 labels per call; fp16 packages pinned by revision and checksum.GLiNER2Manager.classifyConcurrently(text:heads:): several questions per call, and several calls in flight on one model.SortAnythingDemo: 1,000 Wikipedia abstracts sorted into categories chosen at run time (Show and Turbo modes,SORT_AUTOPLAYfor recordings).SortDecisionsDemo: Fastino's Fast Decisions set, every question of a document answered in one call.
Numbers (Apple M5 Pro, 1,000 DBpedia abstracts)
| Core ML (fp16) | gliner2 PyTorch, Mac GPU | |
|---|---|---|
| Time | ~5.9 s | 22.8 s |
| Peak memory | ~1.0 GB | 5.7 GB |
| Weights | 0.92 GB | 1.95 GB (fp32) |
| Accuracy | 89.0% | 89.0% |
Identical predictions on all 1,000. Fast Decisions (public split): 62.9% average, identical to the original model on all 2,900 decisions.
Also since 0.2.0
- GLiClass decision demos and 2048 model comparison (#6); GLiClass models load from Hugging Face (#7).
- GLiNER 2.5 small / base / multilingual Core ML classifiers (#8).
- Sub-1B decision-model runtimes (#9).
- laya runtime and Tetris demo fixes (#5), package docs (#4).
Package
.package(url: "https://github.com/FluidInference/FluidUse.git", from: "0.3.0")