Skip to content

Releases: VinRobotics/vla.cpp

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 13:22
f7e0f7f

vla.cpp now runs 13 VLA architectures, adds Intel (OpenVINO) and Snapdragon (Hexagon, Adreno) backends, and ships portable binaries for Linux x86-64, Linux aarch64 and macOS.

Highlights

  • New models: Octo-Small and TurboVLA.
  • OpenVINO backend for Intel CPUs, iGPUs and NPUs, 3-10x faster than the CPU backend on an Arc iGPU. See docs/backend/ov.md.
  • Snapdragon X on Windows on Arm: Hexagon NPU and Adreno GPU, with CPU fallback. See docs/backend/hexagon-windows.md.
  • Prebuilt binaries for CPU, CUDA 12.8/13.4 (incl. Jetson Orin, Thor, DGX Spark) and Apple Metal, plus a Python wheel. See docs/PREBUILT.md.
  • 6-18% faster inference on CUDA, with identical outputs.
  • More accurate: π0, SmolVLA, GR00T N1.7, VLA-JEPA, Evo-1 and BitVLA now match their reference implementations.
  • Tokenizer in the GGUF: vla-cli --text runs without Python for Octo, π0, π0.5 and OpenVLA-OFT.
  • More stable: the servers no longer crash on bad requests, and the C API is thread-safe.

Breaking changes

  • The x86 CUDA 12.8 tarball is renamed linux-x86_64-cuda-12.8; the macOS tarball ships vla-cli and vla-bench only.
  • The CMake option VLA_OCTO is renamed VLA_SPM (the old name still works).
  • Command-line flags now override --config, and a missing or invalid config file is an error.

Full details in CHANGELOG.md.

What's Changed

  • Support backend OpenVINO by @khanhnd61-vr in #23
  • Octo: ggml inference engine (diffusion + L1/proprio), LIBERO client, and open-loop evaluation by @DuyBaoDOCer in #28
  • Fix GR00T N1.7 relative-action decoding and the ALOHA right-arm client by @hungho77 in #25
  • Feature/add turbovla support by @Zeustakeshi in #29
  • Add Snapdragon X support: Hexagon NPU and Adreno OpenCL with CPU fallback by @khanhnd61-vr in #30
  • vla.cpp performance review by @anindex in #32

New Contributors

Full Changelog: v0.3.0...v0.4.0

v0.3.0

Choose a tag to compare

@khanhnd61-vr khanhnd61-vr released this 24 Aug 04:04
  • add Docker Compose evaluation stack (client image + orchestration + docs) (#15)
  • add end-to-end LIBERO eval on WSL2 (#17)
  • add Apple Silicon (Metal) benchmarks and fix the macOS Metal doc (#20)
  • refactor and get same performance as Pytorch grap capture by using bf16 format (#21)

v0.2.0

Choose a tag to compare

@khanhnd61-vr khanhnd61-vr released this 13 Aug 02:45

vla.cpp v0.2.0

Eleven VLA policies from one self-contained GGUF, now on Intel GPUs as well as CPU, CUDA and Metal.

  • SYCL backend for Intel Arc / Flex / Data Center Max / Xe iGPU; VLA_DEVICE picks the ordinal on CUDA and SYCL alike.
  • Stable C ABI (include/vla.h, libvla) with Python bindings.
  • Four new architectures: π0.5, VLA-Adapter, OpenVLA-OFT, VLA-JEPA.
  • Faster: the compute graph is cached across predict calls in every architecture; Evo-1 encodes all camera views in one pass; BitVLA's ternary GEMM is bank-conflict-free. Opt-in BF16 activations (VLA_*_BF16_ACT) and fused attention (VLA_*_FA) on top.
  • vla-bench, -hf user/repo checkpoint fetch, vla-cli --text.
  • Prebuilt binaries for Linux x86-64 (CPU/CUDA), Linux aarch64 (CPU), macOS Metal, plus a GHCR image.

llama.cpp pinned at b10331. GR00T N1.5 and N1.6 shift by up to 4.6e-4 from an upstream ggml kernel change; the other nine architectures are bit-identical. Full detail in CHANGELOG.md.

v0.1.1

Choose a tag to compare

@anindex anindex released this 06 Jul 11:09
c997aea

First tagged release. C++ inference for 10 VLA policies, validated on the LIBERO object suite (10 tasks).

LIBERO object sweep (RTX 3090, CUDA 12.8)

Model SR client/call (ms)
bitvla 100% 145
evo1 90% 238
gr00t_n1_5 100% 109
gr00t_n1_6 80% 99
gr00t_n1_7 100% 102
openvla_oft 100% 256
pi0 80% 264
pi05 100% 167
smolvla 100% 86
vla_adapter 100% 108

7/10 policies at 100% success. Full report: #14.