Skip to content

SigLIP ViT-B/16 Zero-Shot Classification (FP16)

Choose a tag to compare

@john-rocky john-rocky released this 03 Apr 16:55
· 170 commits to master since this release
7f6e123

Google SigLIP ViT-B/16 converted to 2 CoreML models (FP16).

v2: FP16 replaces INT8. Contrastive models require FP16 for reliable similarity scoring.

Model Size Input Output
SigLIP_ImageEncoder 162 MB 224x224 RGB image L2-normalized 768-dim embedding
SigLIP_TextEncoder 195 MB SentencePiece token IDs L2-normalized 768-dim embedding

Scoring: softmax(image_emb · text_emb * 117.33) across labels.

Total: 357 MB (zipped). iOS 17+.

License: Apache-2.0