Releases: Asher-1/cloudViewer_downloads
Release list
filament
thirdparties
freeimage
mkl_old
yolo_gguf_models
github: https://github.com/Asher-1/ultralytics-ggml
YOLO ggml model card
This directory is the single model store for the C++ integration. PyTorch checkpoints are conversion inputs; GGUF
files are runtime artifacts. Both are ignored by Git and can be regenerated — or downloaded prebuilt: every runtime
GGUF (217 files: the 135 closed-set checkpoints, 13 YOLO-World, 30 YOLOE-26 incl. -pf, 30 yolo26 obb/sem
1024-resolution variants, and the CLIP/MobileCLIP text towers with .ref.npz parity references) is published at
huggingface.co/Asher-1/yolo-gguf and mirrored as GitHub release assets at
cloudViewer_downloads/yolo_gguf_models:
pip install -U "huggingface_hub[cli]"
huggingface-cli download Asher-1/yolo-gguf --local-dir models/gguf
# or the GitHub release mirror
gh release download yolo_gguf_models --repo Asher-1/cloudViewer_downloads --pattern '*.gguf' --dir models/gguf --clobberVerify the 60 obb/sem resolution variants against the tracked checksum list:
cd models/gguf && sha256sum -c SHA256SUMSmodels/
├── MODEL_CARD.md
├── pytorch/ # canonical .pt conversion inputs
│ ├── yolov8{n,s,m,l,x}.pt
│ ├── yolov8{n,s,m,l,x}-seg.pt
│ ├── yolov8{s,m,l,x}-world.pt
│ ├── yoloe-{v8,11}{s,m,l}-seg.pt
│ ├── yoloe-26{n,s,m,l,x}-seg.pt
│ ├── yolo26{n,s,m,l,x}.pt
│ ├── yolo26{n,s,m,l,x}-seg.pt
│ ├── yolo26{n,s,m,l,x}-depth.pt
│ ├── yolo26{n,s,m,l,x}-pose.pt
│ ├── yolo26{n,s,m,l,x}-obb.pt
│ ├── yolo26{n,s,m,l,x}-sem.pt
│ └── yolo26{n,s,m,l,x}-cls.pt
└── gguf/ # generated runtime models
├── <detect-model>-{f32,f16,q8_0}.gguf
├── <detect-model>-seg-{f32,f16,q8_0}.gguf
├── <detect-model>-world-{f32,f16,q8_0}.gguf
├── yoloe-{v8,11}{s,m,l}-seg-{f32,f16,q8_0}.gguf
├── yoloe-26{n,s,m,l,x}-seg-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-depth-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-pose-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-obb-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-obb-1024-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-sem-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-sem-1024-{f32,f16,q8_0}.gguf
├── yolo26{n,s,m,l,x}-cls-{f32,f16,q8_0}.gguf
├── clip-ViT-B-32-{f32,f16,q8_0}.gguf
└── mobileclip2_b-{f32,f16,q8_0}.gguf
Do not put checkpoints in cpp_ggml/ or the repository root. The converter resolves model aliases against
models/pytorch/ and writes to models/gguf/ by default.
Layout migration
Older checkouts may contain the 11 source checkpoints directly under cpp_ggml/ (or a duplicate yolo26n.pt in the
repository root). Those paths are retired. Move any locally retained files once, then remove the old copies:
mkdir -p cpp_ggml/models/pytorch
for name in yolov8n yolov8s yolov8m yolov8l yolov8x yolo26n yolo26s yolo26m yolo26l yolo26x yolo26n-depth; do
test -f "cpp_ggml/$name.pt" && mv "cpp_ggml/$name.pt" "cpp_ggml/models/pytorch/$name.pt"
doneAll conversion, benchmark, parity, and rendering scripts resolve this canonical directory; no script should reference
cpp_ggml/<model>.pt or a root-level checkpoint.
Supported models
| Model | Task | Default input | Recommended use |
|---|---|---|---|
| YOLOv8n | detect | 640 | Lowest detection latency and memory use |
| YOLOv8s | detect | 640 | Small edge deployments needing more capacity than n |
| YOLOv8m | detect | 640 | Balanced accuracy and compute |
| YOLOv8l | detect | 640 | Accuracy-oriented GPU deployment |
| YOLOv8x | detect | 640 | Highest-capacity YOLOv8 integration target |
| YOLO26n | detect | 640 | Lowest-latency end-to-end YOLO26 detector |
| YOLO26s | detect | 640 | Compact end-to-end detector |
| YOLO26m | detect | 640 | Balanced end-to-end detector |
| YOLO26l | detect | 640 | Accuracy-oriented end-to-end detector |
| YOLO26x | detect | 640 | Highest-capacity YOLO26 detection target |
| YOLOv8s-world .. YOLOv8x-world | detect (open-vocabulary) | 640 | Open-vocabulary detection with CLIP text embeddings (--classes, --text-embed) |
YOLOE-v8/11 s..l and YOLOE-26 n..x, -seg |
open-vocabulary instance segment | 640 | Plaintext --classes via native MobileCLIP GGUF, or a YTXT0002 blob |
| YOLOv8n-seg .. YOLOv8x-seg | instance segment | 640 | YOLOv8 boxes + on-device instance masks |
| YOLO26n-seg .. YOLO26x-seg | instance segment | 640 | YOLO26 boxes + on-device instance masks |
| YOLO26n-depth .. YOLO26x-depth | absolute depth | 768 | Monocular metric-depth preview and spatial reasoning |
| YOLO26n-pose .. YOLO26x-pose | keypoints | 640 | COCO-17 person pose (RLE head), boxes + 17 keypoints |
| YOLO26n-obb .. YOLO26x-obb | oriented boxes | 640 / 1024 | DOTA-15 rotated boxes (raw angle, no sigmoid); 640 speed + 1024 native variants |
| YOLO26n-sem .. YOLO26x-sem | semantic seg | 640 / 1024 | Cityscapes-19 dense per-pixel class map; 640 speed + 1024 native variants |
| YOLO26n-cls .. YOLO26x-cls | classification | 224 | ImageNet-1000 logits (checkpoint-baked transforms) |
| CLIP ViT-B/32 | text + image encoder | 224x224 / 77 tokens | 512-d L2-normalised embeddings for semantic similarity search |
Detection models use COCO's 80 classes. YOLO26 detection checkpoints use the end-to-end head exported by the local
Ultralytics checkout. YOLO-World detection models are open-vocabulary: they accept a class list at runtime
(--classes) or a precomputed text embedding blob (--text-embed). The CLIP model (clip-ViT-B-32-f16.gguf by default) is
used to encode class text when --classes is provided without --text-embed; it can also be used independently
for image/text similarity via the similarity subcommand (--model clip-ViT-B-32-f16.gguf --source img.jpg).
Segmentation models additionally emit 32 mask prototypes at one-quarter resolution and compose
instance masks on device. Depth models produce one floating-point distance in meters per source pixel. Pose models
emit one box plus 17 COCO keypoints (x, y, visibility) per person; OBB models emit rotated boxes in the DOTA-15 class
set; semantic models emit an argmax class map on the Cityscapes-19 class set; classify models emit ImageNet-1000
softmax probabilities. The five YOLO26 scales (n/s/m/l/x) share one graph per task; scale changes tensor shapes, not
the public CLI or GGUF contract.
The obb and semantic checkpoints train at 1024, so each scale ships two resolution variants: the canonical file
(yolo26n-obb-f16.gguf) is the 640 speed build and the -1024- file (yolo26n-obb-1024-f16.gguf) is the
checkpoint-native build. The 1024 files are the parity-correct choice — they reproduce Ultralytics Python output
(no end-to-end grid ghost classes, full Cityscapes-19 class set) for about +6% CPU latency and up to ~2.6x compute
on GPU backends. Choose 1024 whenever C++ output must match Python; choose 640 when latency dominates and the
documented 640 grid deviation is acceptable.
YOLOE models consume the raw L2-normalised MobileCLIP feature: the checkpoint's
reprta block is embedded in the GGUF graph (op-graph v4). Pass --classes and
the runtime encodes plaintext end to end with the native MobileCLIP GGUF tower
(mobileclip2_b-{f32,f16,q8_0}.gguf, --text-model); or precompute a checkpoint-agnostic
YTXT0002 blob with scripts/encode_mobileclip_text.py and pass --text-embed.
YOLOv8 detector family
- YOLOv8n is the default for latency-sensitive applications and constrained GPUs.
- YOLOv8s trades a small latency increase for more capacity while remaining suitable for edge deployment.
- YOLOv8m is the balanced choice when throughput and detection quality have similar weight.
- *...
VTK Prebuild Library
vocab_tree
Trellis2 GGUF models
Models 目录指南
github: https://github.com/Asher-1/trellis-ggml
github: https://github.com/Asher-1/RMBG-2.0-GGML
all models available on (some models exceed 2GB): https://huggingface.co/Asher-1/Trellis2-models/tree/main
TRELLIS.2-4B Model Card
This directory holds all model files required for the TRELLIS.2 image-to-3D
pipeline. Two kinds of files live here:
- GGUF runtime models (
*.gguf, top level) — actually loaded by the ggml
inference engine, converted from upstream weights; - Source weights & configs (subdirectories) — upstream HuggingFace
safetensors weights and pipeline configs, used only for conversion/verification,
never loaded at inference time.
All GGUF files are local build artifacts and are excluded from git via
.gitignore (/models/, *.gguf). For benchmarks and parity verification,
see docs/BENCHMARK.md and
docs/VERIFICATION.md.
1. Pipeline Overview
image (RGB/RGBA)
→ [RMBG-2.0] rmbg_f16/f32/q8.gguf Background removal (optional, enabled via --rmbg)
→ [Preprocess] (no model, pure C++)
→ [DINOv3] dino_f16/q8.gguf image → 1029×1024 conditioning tokens
→ [SS-flow DiT] ss_flow_f16/q8.gguf sparse structure flow → z_s (8ch)
→ [SS decoder] ss_dec_f16/q8.gguf 3D-conv → 64³ occupancy → 32³ scaffold
→ [shape-SLAT DiT] slat_flow[_1024]_f16/q8.gguf sparse shape flow (512/1024 res)
→ [shape VAE decoder] shape_dec_f16.gguf sparse ConvNeXt U-Net → dual grid
├→ [meshing] (no model, marching cubes)
└→ [shape VAE enc] shape_enc_f16.gguf reconstructed dual grid → shape SLat
→ [tex-SLAT DiT] tex_slat_flow_512/1024_*.gguf texture flow
→ [texture decoder] tex_dec_f16.gguf sparse 6-channel PBR volume
→ [material sampling] (no model, C++)
The three DiT flows (SS-flow / shape-SLAT / tex-SLAT) share the same 1.3B DiT
architecture (trellis2-slat-flow); they differ only in channel counts and
condition concatenation.
2. Model Inventory
2.1 GGUF runtime models (models/*.gguf)
| Model file | Size | Params | GGUF arch | Source weights |
|---|---|---|---|---|
dino_f16.gguf |
579 MB | 303.1M | trellis2-dino | dinov3-vitl16/ |
dino_q8.gguf |
309 MB | 303.1M | trellis2-dino | same |
rmbg_f16.gguf |
421 MB | 220.7M | rmbg (swin_v1_l) | RMBG-2.0-GGML converter |
rmbg_f32.gguf |
842 MB | 220.7M | rmbg (swin_v1_l) | same |
rmbg_q8.gguf |
247 MB | 220.7M | rmbg (swin_v1_l) | same |
ss_flow_f16.gguf |
2494 MB | 1.29B | trellis2-ss-flow | ss_flow_img_dit_1_3B_64_bf16 |
ss_flow_q8.gguf |
1353 MB | 1.29B | trellis2-ss-flow | same |
ss_dec_f16.gguf |
141 MB | 73.7M | trellis2-ss-dec | ss_dec_conv3d_16l8_fp16 |
ss_dec_q8.gguf |
141 MB | 73.7M | trellis2-ss-dec | same |
slat_flow_f16.gguf |
2494 MB | 1.29B | trellis2-slat-flow | slat_flow_img2shape_dit_1_3B_512_bf16 |
slat_flow_q8.gguf |
1353 MB | 1.29B | trellis2-slat-flow | same |
slat_flow_1024_f16.gguf |
2508 MB | 1.29B | trellis2-slat-flow | slat_flow_img2shape_dit_1_3B_1024_bf16 |
slat_flow_1024_q8.gguf |
1353 MB | 1.29B | trellis2-slat-flow | same |
shape_dec_f16.gguf |
905 MB | 474.2M | trellis2-shape-dec | shape_dec_next_dc_f16c32_fp16 |
shape_enc_f16.gguf |
676 MB | 354.4M | trellis2-shape-enc | shape_enc_next_dc_f16c32_fp16 |
tex_dec_f16.gguf |
905 MB | 474.2M | trellis2-tex-dec | tex_dec_next_dc_f16c32_fp16 |
tex_slat_flow_512_f16.gguf |
2494 MB | 1.29B | trellis2-slat-flow | slat_flow_imgshape2tex_dit_1_3B_512_bf16 |
tex_slat_flow_512_q8.gguf |
1353 MB | 1.29B | trellis2-slat-flow | same |
tex_slat_flow_1024_f16.gguf |
2494 MB | 1.29B | trellis2-slat-flow | slat_flow_imgshape2tex_dit_1_3B_1024_bf16 |
tex_slat_flow_1024_q8.gguf |
1353 MB | 1.29B | trellis2-slat-flow | same |
Params are approximated by summing GGUF tensor shapes. File sizes are the
decimal byte counts fromls -l.
2.2 Source weight directories
| Directory | Content | Size | Purpose |
|---|---|---|---|
TRELLIS.2-4B/ckpts/ |
8 upstream safetensors (5 DiT + 3 VAE) + json | ~16 GB | main pipeline weight source |
TRELLIS.2-4B/pipeline.json |
geometry pipeline config (model map, samplers, normalization) | — | conversion reference |
TRELLIS.2-4B/texturing_pipeline.json |
texturing pipeline config | — | conversion reference |
TRELLIS-image-large/ckpts/ |
ss_dec_conv3d_16l8_fp16.{json,safetensors} |
141 MB | SS decoder weight source |
dinov3-vitl16/ |
DINOv3 ViT-L/16 weights + preprocessor config | 771 MB | image encoder weight source |
3. Per-Model Details
3.1 DINOv3 ViT-L/16 image encoder — dino_f16.gguf / dino_q8.gguf
- Source:
dinov3-vitl16-pretrain-lvd1689m(Meta DINOv3, HF format underdinov3-vitl16/). - Architecture (GGUF
trellis2-dino, 415 tensors): ViT-L/16, hidden=1024,
24 layers, 16 heads, intermediate=4096, patch=16, 4 register tokens, RoPE
(θ=100), no key bias; preprocessing mean=(0.485, 0.456, 0.406),
std=(0.229, 0.224, 0.225). - Role: encodes the preprocessed 512×512 image into
[1, 1029, 1024]
conditioning tokens (1 CLS + 4 register + 1024 patch, taken from the last
layer with affine LN removed) — the visual condition for all three DiT
flows (SS-flow / shape-SLAT / tex-SLAT). - Precision variants: f16 (579 MB) / q8 (309 MB). Q8 quality loss is tiny;
recommended default. - Benchmark (RTX 3060, CUDA): ~1s.
- Note: conditioning tokens are concatenated with F32 tensors downstream,
so token-related weights stay f32.
3.2 RMBG-2.0 background removal — rmbg_f16.gguf / rmbg_f32.gguf / rmbg_q8.gguf
- Source: RMBG-2.0 (BiRefNet family), produced by the
third_party/RMBG-2.0-GGMLconverter. - Architecture (GGUF
rmbg, 742 tensors): backbone = Swin-Transformer-Large
(swin_v1_l), 1024×1024 input, outputs a feathered alpha segmentation map. - Role: removes complex backgrounds so the pipeline reconstructs only the
subject. Not required for solid-color backgrounds; only enabled when
--rmbg MODEL.ggufis passed explicitly — an optional preprocess stage. - Precision variants:
Variant Size CUDA latency Vulkan latency f32 842 MB 644.5 ms 1293.4 ms f16 (recommended) 421 MB 655.2 ms 1278.5 ms q8 247 MB 648.7 ms 1278.5 ms - f16 vs f32 max alpha diff 1.1e-4; f16 is the deployment default.
- Q8 not recommended: saves only 16.9 MiB vs f16, no speedup, and full Q8
exceeds the 2e-3 alpha accuracy gate.
- Note: under heavy load use
--rmbg-device cputo leave VRAM for the main model.
3.3 SS-flow DiT (sparse structure flow) — ss_flow_f16.gguf / ss_flow_q8.gguf
- Source:
ss_flow_img_dit_1_3B_64_bf16(TRELLIS.2-4B main repo). - Architecture (GGUF
trellis2-ss-flow, 640 tensors): resolution=16
(16³=4096 tokens), in/out=8, model_channels=1536, cond_channels=1024,
30 layers, 12 heads, mlp_ratio=5.33, pe_mode=rope, share_mod,
qk_rms_norm (incl. cross). - Role: flow model of the sparse structure stage. 12-step CFG flow-Euler
sampling → sparse structure latentz_s(8ch), which determines the
voxel scaffold of the object. - Benchmark (RTX 3060, CUDA): ~19.8s — matches PyTorch CUDA 19.76s
(diff <0.2%, proving ggml matmul performance parity). - Precision variants: f16 (2494 MB) / q8 (1353 MB). Q8 is essentially
lossless in speed; recommended.
3.4 SS decoder (sparse structure decoder) — ss_dec_f16.gguf / ss_dec_q8.gguf
- Source:
ss_dec_conv3d_16l8_fp16(from the TRELLIS-image-large repo,
stored underTRELLIS-image-large/ckpts/). - Architecture (GGUF
trellis2-ss-dec, 74 tensors): latent_channels=8,
out_channels=1, 3 levels (channels 512/128/32), 2+2 res blocks, layer norm. - Role: decodes
z_sinto 64³ occupancy logits → 32³ voxel scaffold,
the sparse voxel backbone for shape-SLAT. - Benchmark: small model, runs in milliseconds.
- Note: the Q8 variant is the same size as f16 (141 MB) — 3D conv kernels
(ne[0]=3) do not satisfy the ggml alignment constraint and stay at original
precision in practice.
3.5 Shape-SLAT DiT (shape sparse flow) — slat_flow_f16/q8.gguf (512) and slat_flow_1024_f16/q8.gguf (1024)
- Source:
slat_flow_img2shape_dit_1_3B_512_bf16/..._1024_bf16. - Architecture (GGUF
trellis2-slat-flow, 640 tensors): in/out=32,
model_channels=1536, cond_channels=1024, 30 layers, 12 heads, mlp_ratio=5.33,
RoPE, share_mod, qk_rms_norm; embeds shape SLat normalization mean/std
(32 channels).- 512 variant: resolution=32, produces 512³ grids (~1M vertices, default).
- 1024 variant: resolution=64, 1024 cascade, ~5M vertices high-res grids.
- Role: shape flow sampling (12-step CFG) over the 32³ sparse scaffold,
yielding the shape SLat latent (32 channels). - Benchmark (RTX 3060, CUDA): 512 variant sampling ~7s (PyTorch 17.37s,
2.5x faster). - Note: the 1024 HR tokens (~49k) only fit in VRAM via flash attention,
and require DINOv3 encoding at 1024 resolution.
3.6 Shape VAE decoder — shape_dec_f16.gguf
- Source:
shape_dec_next_dc_f16c32_fp16. - Architecture (GGUF
trellis2-shape-dec, 292 tensors): latent_channels=32,
out_channels=7, 5 levels (channels 1024/512/256/128/64, blocks 4/16/8/4/0),
sparse ConvNeXt U-Net, 16× up. - Role: decodes shape SLat into a 7-channel dual grid (occupancy +
features) and outputs subdivision guidance. Performance-critical
bottleneck — ggml CU...
sam_test_data
sam gguf models
Model Zoo — models/
github: https://github.com/Asher-1/sam3-ggml
This directory holds ready-to-run GGUF models for sam3.cpp.
Each file name encodes three things:
<family>_<backbone/size>_<precision>.gguf
| Part | Meaning |
|---|---|
sam3 / sam3-visual |
SAM 3 — ViT-32 backbone + text encoder + DETR detector (850M params) |
sam2 / sam2.1 |
SAM 2 / SAM 2.1 — Meta's Hiera-backbone segmentation models (visual only) |
tiny / small / base_plus / large |
Backbone size (39M / 46M / 81M / 224M params) |
f32 / f16 / q8_0 / q4_1 / q4_0 |
Weight precision (see Precision guide) |
Architecture lineage: this directory covers 2 architectures —
SAM 3 (sam3-*,sam3-visual-*) and the SAM 2 family
(sam2*,sam2.1*). There are no SAM 1 checkpoints (SAM 1 / ViT-B/L/H is a separate
architecture not shipped by this project). The SAM 2 family is visual-only
(points/box + tracking); SAM 3 full adds text-prompted detection (PCS):
type"cat"and get every cat in the image.
Quick pick
| You want… | Pick |
|---|---|
| Text-prompted detection ("type cat, get every cat") | sam3-f16.gguf (1.8 GB) or sam3-q8_0.gguf (1.1 GB) |
| Best visual quality-to-speed balance on GPU | sam2.1_hiera_base_plus_f16.gguf (156 MB) |
| Fastest interactive point/box segmentation on any device | sam2.1_hiera_tiny_q4_0.gguf (23 MB) |
| Best segmentation quality | sam2.1_hiera_large_f16.gguf (431 MB) or _q8_0 (231 MB) |
| Debugging / numerical reference (never for deployment) | sam2.1_hiera_tiny_f32.gguf |
Model files
Sizes below are the actual .gguf files in this directory. Latency is a
single-image PVS run (encode + segment) at 1008×1008 on RTX 3060 CUDA,
point (315,250) on tests/cat.jpg. The current SAM 3 F16 result uses
sam3_encode_image_pvs(), 2 warmups and 7 timed runs (p50); the remaining
rows are the earlier all-model snapshot. score = mask IoU confidence.
SAM 3 (850M params — ViT-32 backbone + text encoder + DETR decoder)
| File | Size | Load | Encode | Segment | Total | score |
|---|---|---|---|---|---|---|
sam3-f32.gguf |
3.3 GB | 3.3 s | 4.2 s | 0.21 s | 7.7 s | 0.953 |
sam3-f16.gguf |
1.8 GB | 0.81 s | 0.566 s | 0.032 s | 1.41 s | 0.953 |
sam3-q8_0.gguf |
1.1 GB | 1.4 s | 3.5 s | 0.20 s | 5.1 s | 0.953 |
sam3-q4_1.gguf |
730 MB | 1.1 s | 3.6 s | 0.21 s | 4.9 s | 0.937 |
sam3-q4_0.gguf |
707 MB | 1.5 s | 3.6 s | 0.25 s | 5.3 s | 0.915 |
Best for: text-prompted detection (PCS) + point/box segmentation (PVS) +
video tracking in one model. The full SAM 3 is the only family here that
supports text prompts; the visual path matches sam3-visual exactly.
SAM 3 Visual (no text encoder — PVS + tracking only)
| File | Size | Load | Encode | Segment | Total | score |
|---|---|---|---|---|---|---|
sam3-visual-f16.gguf |
902 MB | 1.5 s | 2.3 s | 0.22 s | 4.0 s | 0.952 |
sam3-visual-q8_0.gguf |
494 MB | 0.7 s | 2.2 s | 0.22 s | 3.1 s | 0.953 |
sam3-visual-q4_1.gguf |
303 MB | 0.7 s | 2.2 s | 0.21 s | 3.1 s | 0.937 |
sam3-visual-q4_0.gguf |
276 MB | 0.6 s | 2.2 s | 0.23 s | 3.0 s | 0.915 |
Best for: SAM 3-quality segmentation without the text encoder — half the
size and ~40% faster than full SAM 3. Same PVS + tracking capabilities as
sam2.1_hiera_base_plus but with the stronger SAM 3 backbone.
SAM 2 (Hiera backbone, visual only)
| File | Size | Load | Encode | Segment | Total | score |
|---|---|---|---|---|---|---|
sam2_hiera_tiny_f16.gguf |
76 MB | 0.79 s | 1.43 s | 0.37 s | 2.6 s | 0.959 |
sam2_hiera_tiny_f32.gguf |
149 MB | 0.90 s | 0.98 s | 0.20 s | 2.1 s | 0.959 |
sam2_hiera_tiny_q8_0.gguf |
41 MB | 0.39 s | 0.82 s | 0.16 s | 1.4 s | 0.959 |
sam2_hiera_tiny_q4_1.gguf |
25 MB | 0.35 s | 0.81 s | 0.16 s | 1.3 s | 0.930 |
sam2_hiera_tiny_q4_0.gguf |
23 MB | 0.43 s | 0.83 s | 0.17 s | 1.4 s | 0.933 |
sam2_hiera_base_plus_f16.gguf |
156 MB | 0.82 s | 1.22 s | 0.20 s | 2.2 s | 0.957 |
sam2_hiera_base_plus_f32.gguf |
309 MB | 0.97 s | 1.22 s | 0.17 s | 2.4 s | 0.957 |
sam2_hiera_base_plus_q8_0.gguf |
84 MB | 0.93 s | 1.57 s | 0.28 s | 2.8 s | 0.955 |
sam2_hiera_base_plus_q4_1.gguf |
51 MB | 0.73 s | 1.16 s | 0.18 s | 2.1 s | 0.954 |
sam2_hiera_base_plus_q4_0.gguf |
46 MB | 0.65 s | 1.19 s | 0.22 s | 2.1 s | 0.952 |
sam2_hiera_large_f16.gguf |
430 MB | 1.42 s | 1.47 s | 0.25 s | 3.1 s | 0.909 |
SAM 2.1 (improved SAM 2, same Hiera architecture)
| File | Size | Load | Encode | Segment | Total | score |
|---|---|---|---|---|---|---|
sam2.1_hiera_tiny_f32.gguf |
149 MB | 0.78 s | 1.11 s | 0.21 s | 2.1 s | 0.943 |
sam2.1_hiera_tiny_f16.gguf |
76 MB | 0.60 s | 1.03 s | 0.23 s | 1.9 s | 0.943 |
sam2.1_hiera_tiny_q8_0.gguf |
41 MB | 0.62 s | 1.11 s | 0.24 s | 2.0 s | 0.945 |
sam2.1_hiera_tiny_q4_1.gguf |
25 MB | 0.62 s | 0.95 s | 0.19 s | 1.8 s | 0.956 |
sam2.1_hiera_tiny_q4_0.gguf |
23 MB | 0.73 s | 1.13 s | 0.25 s | 2.1 s | 0.927 |
sam2.1_hiera_small_f32.gguf |
176 MB | 0.69 s | 0.88 s | 0.18 s | 1.8 s | 0.945 |
sam2.1_hiera_small_f16.gguf |
90 MB | 0.75 s | 0.99 s | 0.19 s | 1.9 s | 0.945 |
sam2.1_hiera_small_q8_0.gguf |
48 MB | 0.61 s | 0.96 s | 0.19 s | 1.8 s | 0.944 |
sam2.1_hiera_small_q4_1.gguf |
30 MB | 0.72 s | 1.27 s | 0.19 s | 2.2 s | 0.947 |
sam2.1_hiera_small_q4_0.gguf |
27 MB | 0.62 s | 1.11 s | 0.19 s | 1.9 s | 0.949 |
sam2.1_hiera_base_plus_f32.gguf |
309 MB | 0.99 s | 1.25 s | 0.21 s | 2.5 s | 0.953 |
sam2.1_hiera_base_plus_f16.gguf |
156 MB | 0.71 s | 1.09 s | 0.18 s | 2.0 s | 0.953 |
sam2.1_hiera_base_plus_q8_0.gguf |
84 MB | 0.79 s | 1.40 s | 0.21 s | 2.4 s | 0.954 |
sam2.1_hiera_base_plus_q4_1.gguf |
51 MB | 0.74 s | 1.44 s | 0.24 s | 2.4 s | 0.944 |
sam2.1_hiera_base_plus_q4_0.gguf |
46 MB | 0.79 s | 1.34 s | 0.23 s | 2.4 s | 0.936 |
sam2.1_hiera_large_f32.gguf |
857 MB | 1.42 s | 1.57 s | 0.18 s | 3.2 s | 0.940 |
sam2.1_hiera_large_f16.gguf |
431 MB | 1.00 s | 1.29 s | 0.17 s | 2.5 s | 0.940 |
sam2.1_hiera_large_q8_0.gguf |
231 MB | 0.76 s | 1.44 s | 0.19 s | 2.4 s | 0.938 |
sam2.1_hiera_large_q4_1.gguf |
138 MB | 0.75 s | 1.68 s | 0.23 s | 2.7 s | 0.928 |
sam2.1_hiera_large_q4_0.gguf |
124 MB | 0.77 s | 1.49 s | 0.23 s | 2.5 s | 0.900 |
Charts
- Latency chart — 40 locally benchmarked checkpoints (RTX 3060 CUDA)
- Effect grid — 40 locally benchmarked checkpoints on cat.jpg
Precision guide
| Precision | Relative size | Quality | Use |
|---|---|---|---|
f32 |
1.0× | reference | Debugging, numerical checks only — never deploy |
f16 |
0.5× | ≈ f32 | Recommended default — near-lossless, half the size |
q8_0 |
0.25× | very close to f16 | Big models (large/sam3) when f16 is too big |
q4_1 |
~0.14× | good (retains scale + offset) | Aggressive size cuts with better fidelity than q4_0 |
q4_0 |
~0.13× | acceptable for interactive use | Smallest files; quality gap is visible on thin structures |
Size selection guide
| Need | SAM 3 | SAM 3 Visual | base_plus | tiny |
|---|---|---|---|---|
| Text prompts (PCS) | Yes | - | - | - |
| PVS + tracking | Yes | Yes | Yes | Yes |
| Encode latency (RTX 3060) | 0.566 s (F16 PVS) | snapshot: 2.2 s | ~1.1–1.6 s | ~0.8–1.1 s |
| Size (f16) | 1.8 GB | 902 MB | 156 MB | 76 MB |
- SAM 2 vs SAM 2.1: prefer 2.1 for new projects (better training data and
tracking; same architecture, same speed, same sizes). - Video tracking: tiny is the practical choice for interactive playback on
CPU; larger backbones work well on GPU. - Point/box (PVS) + tracking work on every model here; text-prompted
detection (PCS) requires a SAM 3 checkpoint (thesam3-*files above).