Skip to content

v0.4.0

Latest

Choose a tag to compare

@github-actions github-actions released this 30 Sep 13:22
· 16 commits to main since this release
f7e0f7f

vla.cpp now runs 13 VLA architectures, adds Intel (OpenVINO) and Snapdragon (Hexagon, Adreno) backends, and ships portable binaries for Linux x86-64, Linux aarch64 and macOS.

Highlights

  • New models: Octo-Small and TurboVLA.
  • OpenVINO backend for Intel CPUs, iGPUs and NPUs, 3-10x faster than the CPU backend on an Arc iGPU. See docs/backend/ov.md.
  • Snapdragon X on Windows on Arm: Hexagon NPU and Adreno GPU, with CPU fallback. See docs/backend/hexagon-windows.md.
  • Prebuilt binaries for CPU, CUDA 12.8/13.4 (incl. Jetson Orin, Thor, DGX Spark) and Apple Metal, plus a Python wheel. See docs/PREBUILT.md.
  • 6-18% faster inference on CUDA, with identical outputs.
  • More accurate: π0, SmolVLA, GR00T N1.7, VLA-JEPA, Evo-1 and BitVLA now match their reference implementations.
  • Tokenizer in the GGUF: vla-cli --text runs without Python for Octo, π0, π0.5 and OpenVLA-OFT.
  • More stable: the servers no longer crash on bad requests, and the C API is thread-safe.

Breaking changes

  • The x86 CUDA 12.8 tarball is renamed linux-x86_64-cuda-12.8; the macOS tarball ships vla-cli and vla-bench only.
  • The CMake option VLA_OCTO is renamed VLA_SPM (the old name still works).
  • Command-line flags now override --config, and a missing or invalid config file is an error.

Full details in CHANGELOG.md.

What's Changed

  • Support backend OpenVINO by @khanhnd61-vr in #23
  • Octo: ggml inference engine (diffusion + L1/proprio), LIBERO client, and open-loop evaluation by @DuyBaoDOCer in #28
  • Fix GR00T N1.7 relative-action decoding and the ALOHA right-arm client by @hungho77 in #25
  • Feature/add turbovla support by @Zeustakeshi in #29
  • Add Snapdragon X support: Hexagon NPU and Adreno OpenCL with CPU fallback by @khanhnd61-vr in #30
  • vla.cpp performance review by @anindex in #32

New Contributors

Full Changelog: v0.3.0...v0.4.0