Releases: VinRobotics/vla.cpp
Releases · VinRobotics/vla.cpp
Release list
v0.4.0
vla.cpp now runs 13 VLA architectures, adds Intel (OpenVINO) and Snapdragon (Hexagon, Adreno) backends, and ships portable binaries for Linux x86-64, Linux aarch64 and macOS.
Highlights
- New models: Octo-Small and TurboVLA.
- OpenVINO backend for Intel CPUs, iGPUs and NPUs, 3-10x faster than the CPU backend on an Arc iGPU. See docs/backend/ov.md.
- Snapdragon X on Windows on Arm: Hexagon NPU and Adreno GPU, with CPU fallback. See docs/backend/hexagon-windows.md.
- Prebuilt binaries for CPU, CUDA 12.8/13.4 (incl. Jetson Orin, Thor, DGX Spark) and Apple Metal, plus a Python wheel. See docs/PREBUILT.md.
- 6-18% faster inference on CUDA, with identical outputs.
- More accurate: π0, SmolVLA, GR00T N1.7, VLA-JEPA, Evo-1 and BitVLA now match their reference implementations.
- Tokenizer in the GGUF:
vla-cli --textruns without Python for Octo, π0, π0.5 and OpenVLA-OFT. - More stable: the servers no longer crash on bad requests, and the C API is thread-safe.
Breaking changes
- The x86 CUDA 12.8 tarball is renamed
linux-x86_64-cuda-12.8; the macOS tarball shipsvla-cliandvla-benchonly. - The CMake option
VLA_OCTOis renamedVLA_SPM(the old name still works). - Command-line flags now override
--config, and a missing or invalid config file is an error.
Full details in CHANGELOG.md.
What's Changed
- Support backend OpenVINO by @khanhnd61-vr in #23
- Octo: ggml inference engine (diffusion + L1/proprio), LIBERO client, and open-loop evaluation by @DuyBaoDOCer in #28
- Fix GR00T N1.7 relative-action decoding and the ALOHA right-arm client by @hungho77 in #25
- Feature/add turbovla support by @Zeustakeshi in #29
- Add Snapdragon X support: Hexagon NPU and Adreno OpenCL with CPU fallback by @khanhnd61-vr in #30
- vla.cpp performance review by @anindex in #32
New Contributors
- @DuyBaoDOCer made their first contribution in #28
- @hungho77 made their first contribution in #25
- @Zeustakeshi made their first contribution in #29
Full Changelog: v0.3.0...v0.4.0
v0.3.0
v0.2.0
vla.cpp v0.2.0
Eleven VLA policies from one self-contained GGUF, now on Intel GPUs as well as CPU, CUDA and Metal.
- SYCL backend for Intel Arc / Flex / Data Center Max / Xe iGPU;
VLA_DEVICEpicks the ordinal on CUDA and SYCL alike. - Stable C ABI (
include/vla.h,libvla) with Python bindings. - Four new architectures: π0.5, VLA-Adapter, OpenVLA-OFT, VLA-JEPA.
- Faster: the compute graph is cached across
predictcalls in every architecture; Evo-1 encodes all camera views in one pass; BitVLA's ternary GEMM is bank-conflict-free. Opt-in BF16 activations (VLA_*_BF16_ACT) and fused attention (VLA_*_FA) on top. vla-bench,-hf user/repocheckpoint fetch,vla-cli --text.- Prebuilt binaries for Linux x86-64 (CPU/CUDA), Linux aarch64 (CPU), macOS Metal, plus a GHCR image.
llama.cpp pinned at b10331. GR00T N1.5 and N1.6 shift by up to 4.6e-4 from an upstream ggml kernel change; the other nine architectures are bit-identical. Full detail in CHANGELOG.md.
v0.1.1
First tagged release. C++ inference for 10 VLA policies, validated on the LIBERO object suite (10 tasks).
LIBERO object sweep (RTX 3090, CUDA 12.8)
| Model | SR | client/call (ms) |
|---|---|---|
| bitvla | 100% | 145 |
| evo1 | 90% | 238 |
| gr00t_n1_5 | 100% | 109 |
| gr00t_n1_6 | 80% | 99 |
| gr00t_n1_7 | 100% | 102 |
| openvla_oft | 100% | 256 |
| pi0 | 80% | 264 |
| pi05 | 100% | 167 |
| smolvla | 100% | 86 |
| vla_adapter | 100% | 108 |
7/10 policies at 100% success. Full report: #14.