Releases: espetro/dinov2.cpp
Releases · espetro/dinov2.cpp
Release list
v0.5.0
Highlights
- Added a public C API and
libdinov2target, with opaque model/context handles, batch encoding, and model loading from files, buffers, or callbacks. - Added a WebAssembly SIMD128 build and browser demo, now hosted at https://espetro.github.io/dinov2.cpp/.
- Added an optional HTTP embeddings server, plus deduplication, container, and CI visual-regression examples. These remain tier-2, best-effort surfaces.
- Fixed inherited inference and loader issues, including flash-attention padding corruption, GGUF validation, backend device-pointer reads, model-load memory leaks, and positional embedding interpolation.
- Added backbone-only GGUF conversion and publishing for all eight DINOv2 variants, alongside the eight classifier variants. Backbone models support feature extraction and reject classification.
- Added ARM64 benchmark evidence, nightly parity checks, and the JSONL output contract check.
- Improved large-image safety, preprocessing modes, batching, binary embeddings metadata, and clean allocation errors from v0.4.0.
- Added Joaquin Terrasa's copyright notice while retaining the original upstream copyright notice. The project remains MIT licensed.
Compatibility and limitations
- The C API is still unstable during the 0.x series. Pin a release when embedding it.
- The browser demo uses single-threaded WebAssembly SIMD128. Native CPU performance is faster.
- The server and examples are tier-2 and are not production-service guarantees.
- DINOv2 is image-only, not a CLIP-style text-image model. Image decoding is limited to the formats supported by stb_image.
- Benchmark results are hardware-specific; consult
docs/benchmarks.mdanddocs/parity/for methodology and measurements.
Assets
Release binaries are attached below with SHA256 checksums. GGUF model weights are published separately under https://huggingface.co/dinov2-cpp-core.
v0.4.0
Highlights
- Embeddings JSON/JSONL records with a stable contract (
index,image,grid, row-majorpatches), plus a D2EMB v2 preview binary format carrying grid dims. - Batched and multi-image inference via
--batchand repeated-iinputs. - Feature-mode preprocessing is now bounded by default (
--preprocess bounded, shortest edge 518 + true-ceil patch alignment), withhf(HF AutoImageProcessor recipe) andcrop518(fixed 37x37 grid) modes,--no-resizeopt-out, and a--max-tokenscap enforced before graph construction. - Strict numeric parsing for all value flags; graph and model buffer allocation failures exit cleanly instead of aborting.
- Backbone-only DINOv2 GGUF loading and conversion support.
- Parity and benchmark evidence for both small checkpoints (8/8 checks each), plus register-token guidance for feature workflows.
Changelog
[0.4.0] - 2026-09-21
Added
- Return cls, pooled, and score embeddings from dino_predict
- Add output-mode, version, and help flags to CLI parsing
- Emit embeddings JSON and restore top-k output in CLI
- (scripts) Add parity check vs PyTorch reference
- (params) Add n_batch and --batch flag with bounds checking
- (graph) Batch-aware encoder graph in forward_features and attn
- (predict) Multi-image dino_predict API
- (cli) Accept multiple -i inputs and comma-separated image lists
- (cli) Run multi-image inference in n_batch chunks
- (cli) Report per-image throughput in bench output
- (cli) Add preview binary embeddings output
- Harden backbone-only DINOv2 loading
- (preprocess) Bound feature mode with --preprocess, --no-resize, --max-tokens
- (cli) Add index and grid to records, D2EMB v2 header
Fixed
- Preserve aspect ratio in classify preprocess resize
- Pool over actual patch-token count in classify head
- Route dino_model_load logs to stderr
- (scripts) Gate parity on patch flat+mean cosine, keep token min informational
- (dinov2) Exclude register tokens from classify pooling
- (dinov2) Bounds-check value flags in dino_params_parse
- (scripts) Fail parity check on empty image list and CLI timeout
- Add missing scale param to attn declaration
- (cli) Address batched inference review findings
- (cli) Enforce strict numeric parsing
- (cli) Address Task 1 review nits
- (cli) Address binary embeddings review findings
- (ci) Use supported Hugging Face CLI
- (bench) Parse benchmark JSON portably
- Address final parity review findings
- Support backbone-only DINOv2 conversion
- (converter) Harden DINOv2 backbone resolution
- Check ggml graph and model buffer allocation failures
Changed
- (readme) Drop Topics section and topic-count badge
- (plans) Point CUDA smoke-test step at the live HF slug
- Apply clang-format-18 to bench path + bench flag declarations
- Drop unused print_t_f32 debug helper
- Cover classify preprocess aspect and l2_normalize
- Ignore CMake preset build-* directories
- Add cli.md reference and sync CLI docs with embeddings output
- Complete --help blocks and fix model path in cli.md example
- (cli) Fix build preset, threads default, and flag-table lead-in
- Update agent memory for embeddings-dropin learnings
- Batch inference coverage with synthetic tiny model
- Cover flash-attention batch path in dino_predict
- Document batch inference across CLI reference and guides
- (spec) Define register-token recommendations
- Recommend register-token models for feature workflows
- Specify parity benchmark and binary preview
- Add reproducible parity evidence
- (parity) Clarify evidence provenance and downloads
- (bench) Report blocked Ubuntu evidence run
- Record Ubuntu benchmark rerun blocker
- (bench) Record Ubuntu benchmark artifact
- (bench) Preserve model provenance
- (bench) Record current Ubuntu benchmark evidence
- Harden backbone loader review coverage
- (bench) Publish Ubuntu-only benchmark evidence
- Refresh final parity evidence
- Refresh final parity evidence
- Fix regular GGUF checksum
- (spec) Design large-image preprocessing bound and record contract
- Bump project version to 0.4.0
- (spec) Settle preprocess modes, true-ceil alignment, and record contract
- (preprocess) Single resample to bounded target dims, true-ceil ctx sizing
- Record contract, preprocess modes, D2EMB v2
- (parity) Record measured hf-mode residual for tench
- (parity) Record rerun code state and CLI 0.4.0
Miscellaneous
- Add project-scoped mise.toml for clang-format-18 + git-cliff
- Define DINOV2_VERSION from project version
- Add cmake buildPresets so --preset build works
- Merge pull request #10 from espetro/feat/embeddings-dropin
- Merge pull request #11 from espetro/feat/batched-inference
- Merge pull request #12 from espetro/feat/batched-inference
- Merge pull request #13 from espetro/feat/batched-inference
- Merge pull request #14 from espetro/fix/large-image-preprocessing
Full Changelog: v0.3.0...v0.4.0
v0.3.0
v0.2.0
What's Changed
- feat: ggml v0.24.0, HF GGUF publishing pipeline, README rewrite by @espetro in #2
- ci(refactor): port gguf publish to bash script + use ubuntu-latest-large runner by @espetro in #3
- ci(cache): cache .venv-publish + document HF_TOKEN by @espetro in #4
- ci(gguf): fit conversion within ubuntu-latest (~14 GB) via lean deps + cache cleanup by @espetro in #5
- ci(gguf): hardcode HF_HOME (runner.home not allowed in env:) by @espetro in #6
- perf(gguf): relocate caches to /mnt, hf_xet uploader, dynamic matrix, per-run audit logs by @espetro in #7
- fix(ci): emit JSON array from plan job (fromJson requires it, not a space-separated string) by @espetro in #8
Full Changelog: v0.1.0...v0.2.0
v0.1.0
What's Changed
New Contributors
Full Changelog: https://github.com/espetro/dinov2.cpp/commits/v0.1.0