SigLIP ViT-B/16 Zero-Shot Classification (FP16)
Google SigLIP ViT-B/16 converted to 2 CoreML models (FP16).
v2: FP16 replaces INT8. Contrastive models require FP16 for reliable similarity scoring.
| Model | Size | Input | Output |
|---|---|---|---|
| SigLIP_ImageEncoder | 162 MB | 224x224 RGB image | L2-normalized 768-dim embedding |
| SigLIP_TextEncoder | 195 MB | SentencePiece token IDs | L2-normalized 768-dim embedding |
Scoring: softmax(image_emb · text_emb * 117.33) across labels.
Total: 357 MB (zipped). iOS 17+.
License: Apache-2.0