Skip to content

Releases: Asher-1/cloudViewer_downloads

filament

Choose a tag to compare

@Asher-1 Asher-1 released this 22 Jul 03:59
v1.9.9

add test data

thirdparties

Choose a tag to compare

@Asher-1 Asher-1 released this 02 Nov 17:21
1.9.4

add test data

freeimage

Choose a tag to compare

@Asher-1 Asher-1 released this 02 Nov 09:20
1.9.2

add test data

mkl_old

Choose a tag to compare

@Asher-1 Asher-1 released this 02 Nov 09:12
1.9.1

add test data

yolo_gguf_models

Choose a tag to compare

@Asher-1 Asher-1 released this 19 Aug 02:52

github: https://github.com/Asher-1/ultralytics-ggml

YOLO ggml model card

This directory is the single model store for the C++ integration. PyTorch checkpoints are conversion inputs; GGUF
files are runtime artifacts. Both are ignored by Git and can be regenerated — or downloaded prebuilt: every runtime
GGUF (217 files: the 135 closed-set checkpoints, 13 YOLO-World, 30 YOLOE-26 incl. -pf, 30 yolo26 obb/sem
1024-resolution variants, and the CLIP/MobileCLIP text towers with .ref.npz parity references) is published at
huggingface.co/Asher-1/yolo-gguf and mirrored as GitHub release assets at
cloudViewer_downloads/yolo_gguf_models:

pip install -U "huggingface_hub[cli]"
huggingface-cli download Asher-1/yolo-gguf --local-dir models/gguf

# or the GitHub release mirror
gh release download yolo_gguf_models --repo Asher-1/cloudViewer_downloads --pattern '*.gguf' --dir models/gguf --clobber

Verify the 60 obb/sem resolution variants against the tracked checksum list:

cd models/gguf && sha256sum -c SHA256SUMS
models/
├── MODEL_CARD.md
├── pytorch/                    # canonical .pt conversion inputs
│   ├── yolov8{n,s,m,l,x}.pt
│   ├── yolov8{n,s,m,l,x}-seg.pt
│   ├── yolov8{s,m,l,x}-world.pt
│   ├── yoloe-{v8,11}{s,m,l}-seg.pt
│   ├── yoloe-26{n,s,m,l,x}-seg.pt
│   ├── yolo26{n,s,m,l,x}.pt
│   ├── yolo26{n,s,m,l,x}-seg.pt
│   ├── yolo26{n,s,m,l,x}-depth.pt
│   ├── yolo26{n,s,m,l,x}-pose.pt
│   ├── yolo26{n,s,m,l,x}-obb.pt
│   ├── yolo26{n,s,m,l,x}-sem.pt
│   └── yolo26{n,s,m,l,x}-cls.pt
└── gguf/                       # generated runtime models
    ├── <detect-model>-{f32,f16,q8_0}.gguf
    ├── <detect-model>-seg-{f32,f16,q8_0}.gguf
    ├── <detect-model>-world-{f32,f16,q8_0}.gguf
    ├── yoloe-{v8,11}{s,m,l}-seg-{f32,f16,q8_0}.gguf
    ├── yoloe-26{n,s,m,l,x}-seg-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-depth-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-pose-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-obb-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-obb-1024-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-sem-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-sem-1024-{f32,f16,q8_0}.gguf
    ├── yolo26{n,s,m,l,x}-cls-{f32,f16,q8_0}.gguf
    ├── clip-ViT-B-32-{f32,f16,q8_0}.gguf
    └── mobileclip2_b-{f32,f16,q8_0}.gguf

Do not put checkpoints in cpp_ggml/ or the repository root. The converter resolves model aliases against
models/pytorch/ and writes to models/gguf/ by default.

Layout migration

Older checkouts may contain the 11 source checkpoints directly under cpp_ggml/ (or a duplicate yolo26n.pt in the
repository root). Those paths are retired. Move any locally retained files once, then remove the old copies:

mkdir -p cpp_ggml/models/pytorch
for name in yolov8n yolov8s yolov8m yolov8l yolov8x yolo26n yolo26s yolo26m yolo26l yolo26x yolo26n-depth; do
    test -f "cpp_ggml/$name.pt" && mv "cpp_ggml/$name.pt" "cpp_ggml/models/pytorch/$name.pt"
done

All conversion, benchmark, parity, and rendering scripts resolve this canonical directory; no script should reference
cpp_ggml/<model>.pt or a root-level checkpoint.

Supported models

Model Task Default input Recommended use
YOLOv8n detect 640 Lowest detection latency and memory use
YOLOv8s detect 640 Small edge deployments needing more capacity than n
YOLOv8m detect 640 Balanced accuracy and compute
YOLOv8l detect 640 Accuracy-oriented GPU deployment
YOLOv8x detect 640 Highest-capacity YOLOv8 integration target
YOLO26n detect 640 Lowest-latency end-to-end YOLO26 detector
YOLO26s detect 640 Compact end-to-end detector
YOLO26m detect 640 Balanced end-to-end detector
YOLO26l detect 640 Accuracy-oriented end-to-end detector
YOLO26x detect 640 Highest-capacity YOLO26 detection target
YOLOv8s-world .. YOLOv8x-world detect (open-vocabulary) 640 Open-vocabulary detection with CLIP text embeddings (--classes, --text-embed)
YOLOE-v8/11 s..l and YOLOE-26 n..x, -seg open-vocabulary instance segment 640 Plaintext --classes via native MobileCLIP GGUF, or a YTXT0002 blob
YOLOv8n-seg .. YOLOv8x-seg instance segment 640 YOLOv8 boxes + on-device instance masks
YOLO26n-seg .. YOLO26x-seg instance segment 640 YOLO26 boxes + on-device instance masks
YOLO26n-depth .. YOLO26x-depth absolute depth 768 Monocular metric-depth preview and spatial reasoning
YOLO26n-pose .. YOLO26x-pose keypoints 640 COCO-17 person pose (RLE head), boxes + 17 keypoints
YOLO26n-obb .. YOLO26x-obb oriented boxes 640 / 1024 DOTA-15 rotated boxes (raw angle, no sigmoid); 640 speed + 1024 native variants
YOLO26n-sem .. YOLO26x-sem semantic seg 640 / 1024 Cityscapes-19 dense per-pixel class map; 640 speed + 1024 native variants
YOLO26n-cls .. YOLO26x-cls classification 224 ImageNet-1000 logits (checkpoint-baked transforms)
CLIP ViT-B/32 text + image encoder 224x224 / 77 tokens 512-d L2-normalised embeddings for semantic similarity search

Detection models use COCO's 80 classes. YOLO26 detection checkpoints use the end-to-end head exported by the local
Ultralytics checkout. YOLO-World detection models are open-vocabulary: they accept a class list at runtime
(--classes) or a precomputed text embedding blob (--text-embed). The CLIP model (clip-ViT-B-32-f16.gguf by default) is
used to encode class text when --classes is provided without --text-embed; it can also be used independently
for image/text similarity via the similarity subcommand (--model clip-ViT-B-32-f16.gguf --source img.jpg).
Segmentation models additionally emit 32 mask prototypes at one-quarter resolution and compose
instance masks on device. Depth models produce one floating-point distance in meters per source pixel. Pose models
emit one box plus 17 COCO keypoints (x, y, visibility) per person; OBB models emit rotated boxes in the DOTA-15 class
set; semantic models emit an argmax class map on the Cityscapes-19 class set; classify models emit ImageNet-1000
softmax probabilities. The five YOLO26 scales (n/s/m/l/x) share one graph per task; scale changes tensor shapes, not
the public CLI or GGUF contract.

The obb and semantic checkpoints train at 1024, so each scale ships two resolution variants: the canonical file
(yolo26n-obb-f16.gguf) is the 640 speed build and the -1024- file (yolo26n-obb-1024-f16.gguf) is the
checkpoint-native build. The 1024 files are the parity-correct choice — they reproduce Ultralytics Python output
(no end-to-end grid ghost classes, full Cityscapes-19 class set) for about +6% CPU latency and up to ~2.6x compute
on GPU backends. Choose 1024 whenever C++ output must match Python; choose 640 when latency dominates and the
documented 640 grid deviation is acceptable.

YOLOE models consume the raw L2-normalised MobileCLIP feature: the checkpoint's
reprta block is embedded in the GGUF graph (op-graph v4). Pass --classes and
the runtime encodes plaintext end to end with the native MobileCLIP GGUF tower
(mobileclip2_b-{f32,f16,q8_0}.gguf, --text-model); or precompute a checkpoint-agnostic
YTXT0002 blob with scripts/encode_mobileclip_text.py and pass --text-embed.

YOLOv8 detector family

  • YOLOv8n is the default for latency-sensitive applications and constrained GPUs.
  • YOLOv8s trades a small latency increase for more capacity while remaining suitable for edge deployment.
  • YOLOv8m is the balanced choice when throughput and detection quality have similar weight.
  • *...
Read more

VTK Prebuild Library

Choose a tag to compare

@Asher-1 Asher-1 released this 22 Dec 09:18

vtk prebuild library

vocab_tree

Choose a tag to compare

@Asher-1 Asher-1 released this 22 Jan 11:04

publish colmap vocab_tree data

Trellis2 GGUF models

Choose a tag to compare

@Asher-1 Asher-1 released this 05 Aug 08:45

Models 目录指南

github: https://github.com/Asher-1/trellis-ggml
github: https://github.com/Asher-1/RMBG-2.0-GGML
all models available on (some models exceed 2GB): https://huggingface.co/Asher-1/Trellis2-models/tree/main

TRELLIS.2-4B Model Card

This directory holds all model files required for the TRELLIS.2 image-to-3D
pipeline. Two kinds of files live here:

  • GGUF runtime models (*.gguf, top level) — actually loaded by the ggml
    inference engine, converted from upstream weights;
  • Source weights & configs (subdirectories) — upstream HuggingFace
    safetensors weights and pipeline configs, used only for conversion/verification,
    never loaded at inference time.

All GGUF files are local build artifacts and are excluded from git via
.gitignore (/models/, *.gguf). For benchmarks and parity verification,
see docs/BENCHMARK.md and
docs/VERIFICATION.md.


1. Pipeline Overview

image (RGB/RGBA)
  → [RMBG-2.0]           rmbg_f16/f32/q8.gguf            Background removal (optional, enabled via --rmbg)
  → [Preprocess]          (no model, pure C++)
  → [DINOv3]             dino_f16/q8.gguf                image → 1029×1024 conditioning tokens
  → [SS-flow DiT]        ss_flow_f16/q8.gguf             sparse structure flow → z_s (8ch)
  → [SS decoder]         ss_dec_f16/q8.gguf              3D-conv → 64³ occupancy → 32³ scaffold
  → [shape-SLAT DiT]     slat_flow[_1024]_f16/q8.gguf    sparse shape flow (512/1024 res)
  → [shape VAE decoder]  shape_dec_f16.gguf              sparse ConvNeXt U-Net → dual grid
  ├→ [meshing]            (no model, marching cubes)
  └→ [shape VAE enc]     shape_enc_f16.gguf              reconstructed dual grid → shape SLat
     → [tex-SLAT DiT]     tex_slat_flow_512/1024_*.gguf   texture flow
     → [texture decoder] tex_dec_f16.gguf                sparse 6-channel PBR volume
  → [material sampling]   (no model, C++)

The three DiT flows (SS-flow / shape-SLAT / tex-SLAT) share the same 1.3B DiT
architecture (trellis2-slat-flow); they differ only in channel counts and
condition concatenation.


2. Model Inventory

2.1 GGUF runtime models (models/*.gguf)

Model file Size Params GGUF arch Source weights
dino_f16.gguf 579 MB 303.1M trellis2-dino dinov3-vitl16/
dino_q8.gguf 309 MB 303.1M trellis2-dino same
rmbg_f16.gguf 421 MB 220.7M rmbg (swin_v1_l) RMBG-2.0-GGML converter
rmbg_f32.gguf 842 MB 220.7M rmbg (swin_v1_l) same
rmbg_q8.gguf 247 MB 220.7M rmbg (swin_v1_l) same
ss_flow_f16.gguf 2494 MB 1.29B trellis2-ss-flow ss_flow_img_dit_1_3B_64_bf16
ss_flow_q8.gguf 1353 MB 1.29B trellis2-ss-flow same
ss_dec_f16.gguf 141 MB 73.7M trellis2-ss-dec ss_dec_conv3d_16l8_fp16
ss_dec_q8.gguf 141 MB 73.7M trellis2-ss-dec same
slat_flow_f16.gguf 2494 MB 1.29B trellis2-slat-flow slat_flow_img2shape_dit_1_3B_512_bf16
slat_flow_q8.gguf 1353 MB 1.29B trellis2-slat-flow same
slat_flow_1024_f16.gguf 2508 MB 1.29B trellis2-slat-flow slat_flow_img2shape_dit_1_3B_1024_bf16
slat_flow_1024_q8.gguf 1353 MB 1.29B trellis2-slat-flow same
shape_dec_f16.gguf 905 MB 474.2M trellis2-shape-dec shape_dec_next_dc_f16c32_fp16
shape_enc_f16.gguf 676 MB 354.4M trellis2-shape-enc shape_enc_next_dc_f16c32_fp16
tex_dec_f16.gguf 905 MB 474.2M trellis2-tex-dec tex_dec_next_dc_f16c32_fp16
tex_slat_flow_512_f16.gguf 2494 MB 1.29B trellis2-slat-flow slat_flow_imgshape2tex_dit_1_3B_512_bf16
tex_slat_flow_512_q8.gguf 1353 MB 1.29B trellis2-slat-flow same
tex_slat_flow_1024_f16.gguf 2494 MB 1.29B trellis2-slat-flow slat_flow_imgshape2tex_dit_1_3B_1024_bf16
tex_slat_flow_1024_q8.gguf 1353 MB 1.29B trellis2-slat-flow same

Params are approximated by summing GGUF tensor shapes. File sizes are the
decimal byte counts from ls -l.

2.2 Source weight directories

Directory Content Size Purpose
TRELLIS.2-4B/ckpts/ 8 upstream safetensors (5 DiT + 3 VAE) + json ~16 GB main pipeline weight source
TRELLIS.2-4B/pipeline.json geometry pipeline config (model map, samplers, normalization) — conversion reference
TRELLIS.2-4B/texturing_pipeline.json texturing pipeline config — conversion reference
TRELLIS-image-large/ckpts/ ss_dec_conv3d_16l8_fp16.{json,safetensors} 141 MB SS decoder weight source
dinov3-vitl16/ DINOv3 ViT-L/16 weights + preprocessor config 771 MB image encoder weight source

3. Per-Model Details

3.1 DINOv3 ViT-L/16 image encoder — dino_f16.gguf / dino_q8.gguf

  • Source: dinov3-vitl16-pretrain-lvd1689m (Meta DINOv3, HF format under dinov3-vitl16/).
  • Architecture (GGUF trellis2-dino, 415 tensors): ViT-L/16, hidden=1024,
    24 layers, 16 heads, intermediate=4096, patch=16, 4 register tokens, RoPE
    (θ=100), no key bias; preprocessing mean=(0.485, 0.456, 0.406),
    std=(0.229, 0.224, 0.225).
  • Role: encodes the preprocessed 512×512 image into [1, 1029, 1024]
    conditioning tokens (1 CLS + 4 register + 1024 patch, taken from the last
    layer with affine LN removed) — the visual condition for all three DiT
    flows (SS-flow / shape-SLAT / tex-SLAT).
  • Precision variants: f16 (579 MB) / q8 (309 MB). Q8 quality loss is tiny;
    recommended default.
  • Benchmark (RTX 3060, CUDA): ~1s.
  • Note: conditioning tokens are concatenated with F32 tensors downstream,
    so token-related weights stay f32.

3.2 RMBG-2.0 background removal — rmbg_f16.gguf / rmbg_f32.gguf / rmbg_q8.gguf

  • Source: RMBG-2.0 (BiRefNet family), produced by the
    third_party/RMBG-2.0-GGML converter.
  • Architecture (GGUF rmbg, 742 tensors): backbone = Swin-Transformer-Large
    (swin_v1_l), 1024×1024 input, outputs a feathered alpha segmentation map.
  • Role: removes complex backgrounds so the pipeline reconstructs only the
    subject. Not required for solid-color backgrounds; only enabled when
    --rmbg MODEL.gguf is passed explicitly — an optional preprocess stage.
  • Precision variants:
    Variant Size CUDA latency Vulkan latency
    f32 842 MB 644.5 ms 1293.4 ms
    f16 (recommended) 421 MB 655.2 ms 1278.5 ms
    q8 247 MB 648.7 ms 1278.5 ms
    • f16 vs f32 max alpha diff 1.1e-4; f16 is the deployment default.
    • Q8 not recommended: saves only 16.9 MiB vs f16, no speedup, and full Q8
      exceeds the 2e-3 alpha accuracy gate.
  • Note: under heavy load use --rmbg-device cpu to leave VRAM for the main model.

3.3 SS-flow DiT (sparse structure flow) — ss_flow_f16.gguf / ss_flow_q8.gguf

  • Source: ss_flow_img_dit_1_3B_64_bf16 (TRELLIS.2-4B main repo).
  • Architecture (GGUF trellis2-ss-flow, 640 tensors): resolution=16
    (16³=4096 tokens), in/out=8, model_channels=1536, cond_channels=1024,
    30 layers, 12 heads, mlp_ratio=5.33, pe_mode=rope, share_mod,
    qk_rms_norm (incl. cross).
  • Role: flow model of the sparse structure stage. 12-step CFG flow-Euler
    sampling → sparse structure latent z_s (8ch), which determines the
    voxel scaffold of the object.
  • Benchmark (RTX 3060, CUDA): ~19.8s — matches PyTorch CUDA 19.76s
    (diff <0.2%, proving ggml matmul performance parity).
  • Precision variants: f16 (2494 MB) / q8 (1353 MB). Q8 is essentially
    lossless in speed; recommended.

3.4 SS decoder (sparse structure decoder) — ss_dec_f16.gguf / ss_dec_q8.gguf

  • Source: ss_dec_conv3d_16l8_fp16 (from the TRELLIS-image-large repo,
    stored under TRELLIS-image-large/ckpts/).
  • Architecture (GGUF trellis2-ss-dec, 74 tensors): latent_channels=8,
    out_channels=1, 3 levels (channels 512/128/32), 2+2 res blocks, layer norm.
  • Role: decodes z_s into 64³ occupancy logits → 32³ voxel scaffold,
    the sparse voxel backbone for shape-SLAT.
  • Benchmark: small model, runs in milliseconds.
  • Note: the Q8 variant is the same size as f16 (141 MB) — 3D conv kernels
    (ne[0]=3) do not satisfy the ggml alignment constraint and stay at original
    precision in practice.

3.5 Shape-SLAT DiT (shape sparse flow) — slat_flow_f16/q8.gguf (512) and slat_flow_1024_f16/q8.gguf (1024)

  • Source: slat_flow_img2shape_dit_1_3B_512_bf16 / ..._1024_bf16.
  • Architecture (GGUF trellis2-slat-flow, 640 tensors): in/out=32,
    model_channels=1536, cond_channels=1024, 30 layers, 12 heads, mlp_ratio=5.33,
    RoPE, share_mod, qk_rms_norm; embeds shape SLat normalization mean/std
    (32 channels).
    • 512 variant: resolution=32, produces 512³ grids (~1M vertices, default).
    • 1024 variant: resolution=64, 1024 cascade, ~5M vertices high-res grids.
  • Role: shape flow sampling (12-step CFG) over the 32³ sparse scaffold,
    yielding the shape SLat latent (32 channels).
  • Benchmark (RTX 3060, CUDA): 512 variant sampling ~7s (PyTorch 17.37s,
    2.5x faster).
  • Note: the 1024 HR tokens (~49k) only fit in VRAM via flash attention,
    and require DINOv3 encoding at 1024 resolution.

3.6 Shape VAE decoder — shape_dec_f16.gguf

  • Source: shape_dec_next_dc_f16c32_fp16.
  • Architecture (GGUF trellis2-shape-dec, 292 tensors): latent_channels=32,
    out_channels=7, 5 levels (channels 1024/512/256/128/64, blocks 4/16/8/4/0),
    sparse ConvNeXt U-Net, 16× up.
  • Role: decodes shape SLat into a 7-channel dual grid (occupancy +
    features) and outputs subdivision guidance. Performance-critical
    bottleneck
    — ggml CU...
Read more

sam_test_data

Choose a tag to compare

@Asher-1 Asher-1 released this 25 Aug 04:26
add test data

sam gguf models

Choose a tag to compare

@Asher-1 Asher-1 released this 19 Aug 07:03

Model Zoo — models/

github: https://github.com/Asher-1/sam3-ggml

This directory holds ready-to-run GGUF models for sam3.cpp.
Each file name encodes three things:

<family>_<backbone/size>_<precision>.gguf
Part Meaning
sam3 / sam3-visual SAM 3 — ViT-32 backbone + text encoder + DETR detector (850M params)
sam2 / sam2.1 SAM 2 / SAM 2.1 — Meta's Hiera-backbone segmentation models (visual only)
tiny / small / base_plus / large Backbone size (39M / 46M / 81M / 224M params)
f32 / f16 / q8_0 / q4_1 / q4_0 Weight precision (see Precision guide)

Architecture lineage: this directory covers 2 architectures —
SAM 3 (sam3-*, sam3-visual-*) and the SAM 2 family
(sam2*, sam2.1*). There are no SAM 1 checkpoints (SAM 1 / ViT-B/L/H is a separate
architecture not shipped by this project). The SAM 2 family is visual-only
(points/box + tracking); SAM 3 full adds text-prompted detection (PCS):
type "cat" and get every cat in the image.

Quick pick

You want… Pick
Text-prompted detection ("type cat, get every cat") sam3-f16.gguf (1.8 GB) or sam3-q8_0.gguf (1.1 GB)
Best visual quality-to-speed balance on GPU sam2.1_hiera_base_plus_f16.gguf (156 MB)
Fastest interactive point/box segmentation on any device sam2.1_hiera_tiny_q4_0.gguf (23 MB)
Best segmentation quality sam2.1_hiera_large_f16.gguf (431 MB) or _q8_0 (231 MB)
Debugging / numerical reference (never for deployment) sam2.1_hiera_tiny_f32.gguf

Model files

Sizes below are the actual .gguf files in this directory. Latency is a
single-image PVS run (encode + segment) at 1008×1008 on RTX 3060 CUDA,
point (315,250) on tests/cat.jpg. The current SAM 3 F16 result uses
sam3_encode_image_pvs(), 2 warmups and 7 timed runs (p50); the remaining
rows are the earlier all-model snapshot. score = mask IoU confidence.

SAM 3 (850M params — ViT-32 backbone + text encoder + DETR decoder)

File Size Load Encode Segment Total score
sam3-f32.gguf 3.3 GB 3.3 s 4.2 s 0.21 s 7.7 s 0.953
sam3-f16.gguf 1.8 GB 0.81 s 0.566 s 0.032 s 1.41 s 0.953
sam3-q8_0.gguf 1.1 GB 1.4 s 3.5 s 0.20 s 5.1 s 0.953
sam3-q4_1.gguf 730 MB 1.1 s 3.6 s 0.21 s 4.9 s 0.937
sam3-q4_0.gguf 707 MB 1.5 s 3.6 s 0.25 s 5.3 s 0.915

Best for: text-prompted detection (PCS) + point/box segmentation (PVS) +
video tracking in one model. The full SAM 3 is the only family here that
supports text prompts; the visual path matches sam3-visual exactly.

SAM 3 Visual (no text encoder — PVS + tracking only)

File Size Load Encode Segment Total score
sam3-visual-f16.gguf 902 MB 1.5 s 2.3 s 0.22 s 4.0 s 0.952
sam3-visual-q8_0.gguf 494 MB 0.7 s 2.2 s 0.22 s 3.1 s 0.953
sam3-visual-q4_1.gguf 303 MB 0.7 s 2.2 s 0.21 s 3.1 s 0.937
sam3-visual-q4_0.gguf 276 MB 0.6 s 2.2 s 0.23 s 3.0 s 0.915

Best for: SAM 3-quality segmentation without the text encoder — half the
size and ~40% faster than full SAM 3. Same PVS + tracking capabilities as
sam2.1_hiera_base_plus but with the stronger SAM 3 backbone.

SAM 2 (Hiera backbone, visual only)

File Size Load Encode Segment Total score
sam2_hiera_tiny_f16.gguf 76 MB 0.79 s 1.43 s 0.37 s 2.6 s 0.959
sam2_hiera_tiny_f32.gguf 149 MB 0.90 s 0.98 s 0.20 s 2.1 s 0.959
sam2_hiera_tiny_q8_0.gguf 41 MB 0.39 s 0.82 s 0.16 s 1.4 s 0.959
sam2_hiera_tiny_q4_1.gguf 25 MB 0.35 s 0.81 s 0.16 s 1.3 s 0.930
sam2_hiera_tiny_q4_0.gguf 23 MB 0.43 s 0.83 s 0.17 s 1.4 s 0.933
sam2_hiera_base_plus_f16.gguf 156 MB 0.82 s 1.22 s 0.20 s 2.2 s 0.957
sam2_hiera_base_plus_f32.gguf 309 MB 0.97 s 1.22 s 0.17 s 2.4 s 0.957
sam2_hiera_base_plus_q8_0.gguf 84 MB 0.93 s 1.57 s 0.28 s 2.8 s 0.955
sam2_hiera_base_plus_q4_1.gguf 51 MB 0.73 s 1.16 s 0.18 s 2.1 s 0.954
sam2_hiera_base_plus_q4_0.gguf 46 MB 0.65 s 1.19 s 0.22 s 2.1 s 0.952
sam2_hiera_large_f16.gguf 430 MB 1.42 s 1.47 s 0.25 s 3.1 s 0.909

SAM 2.1 (improved SAM 2, same Hiera architecture)

File Size Load Encode Segment Total score
sam2.1_hiera_tiny_f32.gguf 149 MB 0.78 s 1.11 s 0.21 s 2.1 s 0.943
sam2.1_hiera_tiny_f16.gguf 76 MB 0.60 s 1.03 s 0.23 s 1.9 s 0.943
sam2.1_hiera_tiny_q8_0.gguf 41 MB 0.62 s 1.11 s 0.24 s 2.0 s 0.945
sam2.1_hiera_tiny_q4_1.gguf 25 MB 0.62 s 0.95 s 0.19 s 1.8 s 0.956
sam2.1_hiera_tiny_q4_0.gguf 23 MB 0.73 s 1.13 s 0.25 s 2.1 s 0.927
sam2.1_hiera_small_f32.gguf 176 MB 0.69 s 0.88 s 0.18 s 1.8 s 0.945
sam2.1_hiera_small_f16.gguf 90 MB 0.75 s 0.99 s 0.19 s 1.9 s 0.945
sam2.1_hiera_small_q8_0.gguf 48 MB 0.61 s 0.96 s 0.19 s 1.8 s 0.944
sam2.1_hiera_small_q4_1.gguf 30 MB 0.72 s 1.27 s 0.19 s 2.2 s 0.947
sam2.1_hiera_small_q4_0.gguf 27 MB 0.62 s 1.11 s 0.19 s 1.9 s 0.949
sam2.1_hiera_base_plus_f32.gguf 309 MB 0.99 s 1.25 s 0.21 s 2.5 s 0.953
sam2.1_hiera_base_plus_f16.gguf 156 MB 0.71 s 1.09 s 0.18 s 2.0 s 0.953
sam2.1_hiera_base_plus_q8_0.gguf 84 MB 0.79 s 1.40 s 0.21 s 2.4 s 0.954
sam2.1_hiera_base_plus_q4_1.gguf 51 MB 0.74 s 1.44 s 0.24 s 2.4 s 0.944
sam2.1_hiera_base_plus_q4_0.gguf 46 MB 0.79 s 1.34 s 0.23 s 2.4 s 0.936
sam2.1_hiera_large_f32.gguf 857 MB 1.42 s 1.57 s 0.18 s 3.2 s 0.940
sam2.1_hiera_large_f16.gguf 431 MB 1.00 s 1.29 s 0.17 s 2.5 s 0.940
sam2.1_hiera_large_q8_0.gguf 231 MB 0.76 s 1.44 s 0.19 s 2.4 s 0.938
sam2.1_hiera_large_q4_1.gguf 138 MB 0.75 s 1.68 s 0.23 s 2.7 s 0.928
sam2.1_hiera_large_q4_0.gguf 124 MB 0.77 s 1.49 s 0.23 s 2.5 s 0.900

Charts

Precision guide

Precision Relative size Quality Use
f32 1.0× reference Debugging, numerical checks only — never deploy
f16 0.5× ≈ f32 Recommended default — near-lossless, half the size
q8_0 0.25× very close to f16 Big models (large/sam3) when f16 is too big
q4_1 ~0.14× good (retains scale + offset) Aggressive size cuts with better fidelity than q4_0
q4_0 ~0.13× acceptable for interactive use Smallest files; quality gap is visible on thin structures

Size selection guide

Need SAM 3 SAM 3 Visual base_plus tiny
Text prompts (PCS) Yes - - -
PVS + tracking Yes Yes Yes Yes
Encode latency (RTX 3060) 0.566 s (F16 PVS) snapshot: 2.2 s ~1.1–1.6 s ~0.8–1.1 s
Size (f16) 1.8 GB 902 MB 156 MB 76 MB
  • SAM 2 vs SAM 2.1: prefer 2.1 for new projects (better training data and
    tracking; same architecture, same speed, same sizes).
  • Video tracking: tiny is the practical choice for interactive playback on
    CPU; larger backbones work well on GPU.
  • Point/box (PVS) + tracking work on every model here; text-prompted
    detection (PCS) requires a SAM 3 checkpoint
    (the sam3-* files above).