Skip to content

v0.2.0

Choose a tag to compare

@khanhnd61-vr khanhnd61-vr released this 13 Aug 02:45
· 30 commits to main since this release

vla.cpp v0.2.0

Eleven VLA policies from one self-contained GGUF, now on Intel GPUs as well as CPU, CUDA and Metal.

  • SYCL backend for Intel Arc / Flex / Data Center Max / Xe iGPU; VLA_DEVICE picks the ordinal on CUDA and SYCL alike.
  • Stable C ABI (include/vla.h, libvla) with Python bindings.
  • Four new architectures: π0.5, VLA-Adapter, OpenVLA-OFT, VLA-JEPA.
  • Faster: the compute graph is cached across predict calls in every architecture; Evo-1 encodes all camera views in one pass; BitVLA's ternary GEMM is bank-conflict-free. Opt-in BF16 activations (VLA_*_BF16_ACT) and fused attention (VLA_*_FA) on top.
  • vla-bench, -hf user/repo checkpoint fetch, vla-cli --text.
  • Prebuilt binaries for Linux x86-64 (CPU/CUDA), Linux aarch64 (CPU), macOS Metal, plus a GHCR image.

llama.cpp pinned at b10331. GR00T N1.5 and N1.6 shift by up to 4.6e-4 from an upstream ggml kernel change; the other nine architectures are bit-identical. Full detail in CHANGELOG.md.