Repository navigation
vla.cpp now runs 13 VLA architectures, adds Intel (OpenVINO) and Snapdragon (Hexagon, Adreno) backends, and ships portable binaries for Linux x86-64, Linux aarch64 and macOS.
Highlights
- New models: Octo-Small and TurboVLA.
- OpenVINO backend for Intel CPUs, iGPUs and NPUs, 3-10x faster than the CPU backend on an Arc iGPU. See docs/backend/ov.md.
- Snapdragon X on Windows on Arm: Hexagon NPU and Adreno GPU, with CPU fallback. See docs/backend/hexagon-windows.md.
- Prebuilt binaries for CPU, CUDA 12.8/13.4 (incl. Jetson Orin, Thor, DGX Spark) and Apple Metal, plus a Python wheel. See docs/PREBUILT.md.
- 6-18% faster inference on CUDA, with identical outputs.
- More accurate: π0, SmolVLA, GR00T N1.7, VLA-JEPA, Evo-1 and BitVLA now match their reference implementations.
- Tokenizer in the GGUF:
vla-cli --textruns without Python for Octo, π0, π0.5 and OpenVLA-OFT. - More stable: the servers no longer crash on bad requests, and the C API is thread-safe.
Breaking changes
- The x86 CUDA 12.8 tarball is renamed
linux-x86_64-cuda-12.8; the macOS tarball shipsvla-cliandvla-benchonly. - The CMake option
VLA_OCTOis renamedVLA_SPM(the old name still works). - Command-line flags now override
--config, and a missing or invalid config file is an error.
Full details in CHANGELOG.md.
What's Changed
- Support backend OpenVINO by @khanhnd61-vr in #23
- Octo: ggml inference engine (diffusion + L1/proprio), LIBERO client, and open-loop evaluation by @DuyBaoDOCer in #28
- Fix GR00T N1.7 relative-action decoding and the ALOHA right-arm client by @hungho77 in #25
- Feature/add turbovla support by @Zeustakeshi in #29
- Add Snapdragon X support: Hexagon NPU and Adreno OpenCL with CPU fallback by @khanhnd61-vr in #30
- vla.cpp performance review by @anindex in #32
New Contributors
- @DuyBaoDOCer made their first contribution in #28
- @hungho77 made their first contribution in #25
- @Zeustakeshi made their first contribution in #29
Full Changelog: v0.3.0...v0.4.0