Skip to content

History / Review Discrete Diffusion VLA

Revisions

  • Source-verification audit: fix fabrications, rename PI slugs, link hygiene, add foundational pages - Re-verified all 195 pages vs original sources; removed 47 confirmed fabricated tables/numbers, restored 19 false-positive deletions (full-PDF re-check) - Renamed PI tech-report slugs ICLR-2026-pi07/pi06/RECAP -> PI-pi07/PI-pi06/PI-RECAP (these are PI technical reports, not ICLR 2026 papers); updated 171 wikilinks - Fixed 60 broken wikilinks -> 0 dangling across the wiki - Added foundational pages: OpenVLA, ReKep, AgiBot World Colosseo, RoboBrain 2.0; linked from Home + sidebar + venue indexes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Jun 1, 2026
  • Add in-depth review for Discrete Diffusion VLA (ICLR 2026) Review-Discrete-Diffusion-VLA.md is a long-form companion to the existing ICLR 2026 DDVLA summary page. Sources are arXiv 2508.20072v3 and the official GitHub release. Key positioning: DDVLA is the ICLR 2026 flagship for the "unified token stream" architectural family (Category D) - the direct counter- proposal to the pi-series' separate flow-matching action expert. Structure: - Architectural context table: DDVLA vs pi-series (same-stack MoE) vs Fast-in-Slow (embedded) - 3 competing philosophies at ICLR 2026 - Figure 1 (paradigm comparison: Cont-Diffusion / AR / BERT / Discrete-Diffusion-with-remasking) + Figure 2 (architecture) embedded from the paper with attribution - Full method: * 256-bin quantile tokenization, 7 tokens/timestep * Prismatic-7B backbone (SigLIP+DINOv2 ViT + Llama 2) * Bidirectional attention over action tokens * Masked cross-entropy training objective (no separate flow loss) * Adaptive decoding with max-confidence OR confidence-gap selection * Secondary re-masking via threshold + residual-drop checks * T=12 refinement rounds; linear temperature decay - All accuracy numbers: * LIBERO 96.3% avg (Spatial 97.2, Object 98.6, Goal 97.4, Long 92.0) * SimplerEnv-Fractal 71.2% visual matching (vs pi0 58.8, pi0-FAST 61.9); overall 64.1% * SimplerEnv-Bridge 54.2% (+6.4 over pi0-FAST, +14.1 over pi0) - Ablations: * Adaptive ordering progression: 95.6 -> 95.8 -> 96.6 -> 97.0 -> 97.4 * Temperature schedule: argmax 96.2, fixed 96.4, linear decay 97.4 * OOD robustness: -1.4% language-aug (vs -8.0% parallel one-shot, -2.4% continuous diffusion); -21% vision-aug (best among tested) - Inference efficiency: 12 NFEs vs AR's 56 = 4.7x fewer; 68.8 ms on H800 = 2x AR speedup, matches continuous-diffusion latency - Mechanistic argument: parallel + iterative + adaptive order + re-masking + single CE loss - no other decoder family combines all - Limitations: no head-to-head vs pi0.6/pi0.7; LIBERO+SimplerEnv only (no real-robot); fixed chunk length; 256-bin quantization ceiling; small (+0.9%) accuracy gain over same-tokenization OpenVLA-OFT Assets: - assets/ddvla_fig1_paradigm.png (Fig 1 paradigm comparison) - assets/ddvla_fig2_architecture.png (Fig 2 full architecture) Navigation: - ICLR-2026-Discrete-Diffusion-VLA.md adds "In-depth review" link - _Sidebar.md per-paper reviews section adds DDVLA (5th per-paper) - Home.md per-paper reviews table + What's-new surface the addition Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Apr 20, 2026