-
Notifications
You must be signed in to change notification settings - Fork 0
Home
A living research wiki for Vision-Language-Action models Β· robot manipulation Β· humanoid intelligence Β· tactile & force interaction Β· world models.
| π 396 pages | π 8 venues | π 16+ cross-paper reviews | π updated 2026-07-25 |
|---|
New here? Start with one of these: π§ VLA Architectures β choose the model family & action decoder Β· π Ο Series Evolution β the production-VLA baseline lineage Β· π IROS 2026 Survey π Β· RSS 2026 Survey β the latest venue surveys Β· π Latest Papers β preprint tracker (pre-publication reviews) Β· πΈοΈ Knowledge Graph β the visual map
Pick your question β open the fold-out for a 4-line orientation β follow the π deep-dive link for the full survey (trend arc Β· approach taxonomy Β· limitations). Every deep-dive page carries a matching "State of the Field" section, updated Aug 2026.
Which VLA architecture should we use? β flow matching dominates; ordered tokens revived AR
π Deep dive β VLA Architectures
| π Trend | AR tokens (2024) β flow-matching experts standard (2025) β flow universal in flagships + AR revived via OAT's anytime prefix decoding (2026) |
| βοΈ Approaches | AR discrete tokens (LLM-native / lossy, slow) Β· flow-matching expert (precise, production-proven / no native log-probs) Β· discrete diffusion in-VLM (parallel decode / little traction) |
| No matched-scale comparison across the three families; latency rarely published |
How should the VLM connect to the action expert? β evidence tilts to cross-attention, margins small
π Deep dive β VLMβAction Connection
| π Trend | 2025 four-way standoff β first direct ablations in 2026: cross-attention wins in RobotManip (87.5 vs 87.0) and MM-DiT beats naive DiT in Ξ¨β |
| βοΈ Approaches | same-stack MoE+prefix-KV (tight / hard to retrofit) Β· cross-attention (decoupled, small experts / one-way) Β· concatenation (simplest / re-attends per step) Β· adaLN token (cheapest / bottleneck) |
| ~0.5 pp margins on one benchmark; the two Qwen flagships disagree internally |
How should attention / KV-cache be designed? β the runtime now drives the choice
π Deep dive β VLA Attention
| π Trend | Design follows the runtime since 2025 (RTC/streaming/prefix-KV); 2026 adds linear-attention backbones (Qwen3.5) with zero control-quality ablation |
| βοΈ Approaches | prefix-KV + expert branch (cache reuse) Β· full re-attention per chunk (freshness) Β· streaming-token designs |
| Latency is the field's least-reported number; multi-view KV behavior is folklore |
Should vision be built outside the VLM? β the ViT is the proven bottleneck; fix = action supervision
π Deep dive β Independent Visual Representation
| π Trend | VLM4VLA: VLM scores don't predict control; frozen ViT β21~42 pp; action-supervised vision FT +18.1 pp β the one proven fix. 2026 splits into geometry add-ons vs dynamics encoders |
| βοΈ Approaches | action-align the VLM ViT (cheap, proven) Β· bolt-on metric 3D (spatial gains / calibration tax) Β· video-dynamics encoders (physics priors / loses VLM semantics) |
| Effects visible only under OOD stress; no standardized recipe |
Which training framework / codebase? β modular vs pipeline; RL support is the new differentiator
π Deep dive β VLA Training Frameworks
| π Trend | StarVLA (modular) vs TRI VLA Foundry (pipeline) β RSS 2026 adds the RL tier (RLux-VLA) |
| Nothing covers pretrainβSFTβRLβevaluation end-to-end; flagship recipes irreproducible in public stacks |
How do we make inference real-time? β from test-time patches to trained-in continuation
π Deep dive β Real-Time Execution
| π Trend | RTC (test-time, 2025) β training-time RTC β Legato native continuation (beats RTC ~10%) + OAT anytime decoding (2026) |
| βοΈ Approaches | test-time guidance (drop-in / unstable on some models) Β· trained continuation (smooth / retrain) Β· anytime tokens (compute dial / AR-only) Β· async dual-rate (decoupled rates / complexity) |
| ms/chunk almost never published; robustness features raise the denoise-step bill (4β10) |
How does a deployed policy improve from experience? π β production-proven at RSS 2026
π Deep dive β RL for VLA
| π Trend | "Impractical" (2025) β log-prob wall cracked 3 ways β Ο*0.6/RECAP in production: 2Γ throughput, ~Β½ failures on laundry/boxes/espresso |
| βοΈ Approaches | advantage conditioning (no log-probs, eats corrections / coarse credit) Β· direct PG on flow (principled / machinery) Β· residual RL (safe / capped) Β· world-model RFT (no rollouts / WM fidelity) |
| Loops are domain-narrow; exploration safety procedural; forgetting is milder than feared (ICML Oral: pretrained VLAs resist it) but untested under repeated RL |
How do we handle long-horizon tasks? β from bigger context to agentic decomposition
π Deep dive β VLA Memory Β· System 0/1/2
| π Trend | In-policy memory modules (2025) β selective key-frame history + agentic planner/executor with notebook memory (RobotNav: EQA SOTA, β77% steps) |
| βοΈ Approaches | in-policy memory (fast / shortcut risk) Β· episodic retrieval (scales / retrieval-bound) Β· agentic notebook (auditable / planner latency) |
| Manipulation autonomy record is 11 stages / 2.5 min; no agentic-manipulation demo yet |
Do world models help VLA? β no longer hypothetical: dynamics pre-training beats VLAs on hardware
π Deep dive β World Models Β· WAM vs VLA
| π Trend | Serving role (2025) β challenger evidence at RSS 2026: LDA-1B +48% dexterous over Ο0.5; mimic-video 10Γ sample efficiency |
| βοΈ Approaches | data engine (mature / fidelity ceiling) Β· evaluator (promising / action-input gap β ICML's dWorldEval is the first crack) Β· VAM backbone (dynamics priors / loses VLM depth) Β· unified WM+policy (all data tiers / heaviest) Β· aux losses (free / weak) |
| Contact physics fidelity; VAM claims await matched-scale replication; 20B video models price out labs |
Can human video substitute for robot data? π β yes; the field split on how
π Deep dive β Human Video β Robot Transfer
| π Trend | Retargeting pipelines (2024β25) β three camps at RSS 2026: emergence (PI: co-train, ~2Γ above a diversity threshold) vs decoupling (Ξ¨β: 800 h + 30 h beats 10Γ corpora) vs synthesis (RobotManip: 24.8k h H2R) |
| βοΈ Approaches | co-train (no pipeline / threshold, gripper-only evidence) Β· staged decouple (data-efficient, humanoid-proven / 2-stage) Β· synthesize (unlimited scale / artifact ceiling) |
| Nobody has run both recipes on the same platform β the field's most valuable missing experiment |
What data should we co-train on? π β the most settled question on this map
π Deep dive β LBM Co-training (+ RSS 89-policy sequel)
| π Trend | Folklore β measurement: two LBM studies (89 policies, 58k+2,835 rollouts) + independent Qwen/VLM4VLA confirmations |
| β Verdicts | VL + cross-embodiment data help cumulatively Β· discrete robot-action tokens don't (3Γ replicated; latent-action tokens as VLM supervision do β ICML Oral) Β· co-training must be joint, not sequential |
| Mixture ratios are art (Ξ»=0.1 vs 1.0, undiscussed); no per-sample attribution |
How do we evaluate policies credibly? π β the in-distribution era is ending
π Deep dive β VLA Evaluation
| π Trend | Triple indictment (from-scratch β pretrained in-distribution; LIBERO-X 39.4β8.2 collapse; weak sim-real correlation) β RSS 2026 infrastructure: PolaRiS real-to-sim with validated rank correlation |
| βοΈ Approaches | static sim (cheap / saturated) Β· perturbation pyramids (diagnostic / still sim) Β· real-to-sim (reality-anchored / contact physics) Β· rigorous real stats (ground truth / cost) |
| Contact-rich sim evaluation missing everywhere; OOD reporting still voluntary |
How do we improve dexterous manipulation? β touch moved from observation to prediction
π Deep dive β Dexterous Manipulation Β· Tactile VLA
| π Trend | RL still owns in-hand skills; IL changed paradigm at RSS 2026 β ViTacFormer forecasts future contact (+50%, 11-stage record); CGP commands contacts, not poses |
| βοΈ Approaches | sim-to-real RL (reactive / per-skill, tactile sim gap) Β· predictive visuo-tactile IL (breadth / rig cost) Β· dexterous VLAs (language / trails specialists) |
| No dexterous foundation model; precision insertion unsolved everywhere (Ξ¨β 2/10, screws 0/10) |
How do we generalize across robots? β alignment is the precondition for data scaling
π Deep dive β Cross-Embodiment
| π Trend | Question flipped in 2026: RobotManip showed misaligned action spaces produce no scaling law; camera-frame EEF scales + transfers zero-shot. RSS extended to hands (DexGrasp-Zero 85% zero-shot, One-Hand 81.9%) |
| βοΈ Approaches | canonical padded tensors (simple, proven / curated slots) Β· camera-frame delta EEF (best transfer / calibration tax) Β· morphology graphs/URDFs (anatomy-grounded / hands-only) Β· prompts + history (no arch change / weak alone) |
| Joint-space zero-shot <5%; Franka-class morphology gaps resist; gripperβhand untried |
How do we run a VLA on a humanoid? β the triple-system recipe became the open reference
π Deep dive β Humanoid VLA Β· System 0/1/2
| π Trend | RSS 2026 = humanoid loco-manipulation's coming-out; Ξ¨β's VLM + MM-DiT + RL-lower-body triple system beats 10Γ-data baselines by +40 pp |
| βοΈ Approaches | whole-body end-to-end (expressive / unstable, data-hungry) Β· triple-system (stable, data-efficient / agility capped) Β· teleop rigs vs robot-free human interfaces (quality vs scale) |
| Per-task fine-tuning in every loop; no humanoid OOD protocol β evaluation lags arms by a generation |
Newest first β full history in Changelog. Pre-publication preprint reviews live in Latest Papers.
| Date | Page | What it adds |
|---|---|---|
| 07-25 | RSS 2026 survey + 15 paper pages | Sydney; 210 papers, ~116 in scope; Ο*0.6/RECAP flagship; figure-illustrated reviews |
| 07-25 | Ξ¨β (in-depth) | Open humanoid foundation model β 800 h human video + 30 h robot data beats 10Γ corpora |
| 07-24 | Qwen Team's VLA Program | Cross-paper: VLM4VLA β Qwen-VLA β Qwen-Robot Suite (5 in-depth reviews) |
| 07-24 | Qwen-RobotManip Β· RobotNav Β· RobotWorld | The Qwen-Robot Suite, deep-read |
| 06-11 | VLA Training Frameworks | StarVLA vs TRI VLA Foundry |
| 06-10 | RoboMME (in-depth) | Memory-implementation breakdown |
β Topic reviews catalog β cross-paper topic reviews, lab programs, latest-paper reviews Β· Per-paper long-forms β single-paper deep-dives.
| Most-used entries | |
|---|---|
| Topic reviews | VLA Architectures Β· VLA Hybrid Architectures π Β· Multi-Task VLA π Β· VLMβAction Β· RL for VLA Β· World Models Β· Dexterous Β· Dex-Hand Data Pyramid π Β· Tactile Β· Cross-Embodiment Β· Single-Checkpoint Multi-Robot π Β· Egocentric Video Pre-Training π Β· Humanoid Β· Memory Β· In-Context Imitation π |
| Lab programs | Ο series Β· GR00T N1βN1.7 Β· RoboTTT (context scaling) π Β· Qwen VLA program Β· Ξ¨β Β· DreamZero π |
| ML foundations | ML hub Β· Attention Variants Β· Normalization |
| Venue | Year | Survey / Index | Character |
|---|---|---|---|
| RSS | 2026 | Survey π | Method frontier β the improvement loop, humanoids, hands, evaluation (210 papers) |
| ICRA | 2026 | Survey | Deployment & sensors at scale (728 in-scope papers, 12 topic pages) |
| ICML | 2026 | Index | 99 manipulation papers + π Top-20 |
| CVPR | 2026 | Survey | Vision-first VLA / perception (also 2025) |
| ICLR | 2026 | Survey | Architecture & representation (~211 papers) |
| NeurIPS | 2025 | Survey | Foundation-model training methods |
| CoRL | 2026 | Survey π | Robot learning (687 accepted; preliminary β official list pending) |
| CoRL | 2025 | Survey | Robot learning / benchmarks |
| IROS | 2026 | Survey π | Robustness/efficiency/deployment of learned & VLA policies (1,900+ papers; titles-level, pre-conf) |
| IROS | 2025 | Survey | Hardware, systems, deployment |
OpenVLA (open 7B AR baseline) Β· ReKep (training-free keypoint constraints) Β· AgiBot World Colosseo (1M+ trajectory dataset) Β· RoboBrain 2.0 (embodied-reasoning VLM)
π Contributing: page templates & house rules β Maintenance. Summaries are compiled from public sources and are not substitutes for the original papers β verify numbers before citing.
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)