-
Notifications
You must be signed in to change notification settings - Fork 0
IROS 2026 3D FlowMatch Actor
Venue: IROS 2026 (Pittsburgh) Β· paper #4030 Β· Carnegie Mellon University Β· NVIDIA Β· National Taiwan University (Gkanatsios, Xu, Bronars, Mousavian, Ke, Fragkiadaki). Paper: arXiv 2508.11002 Β· project Β· code. The bimanual-SOTA datapoint of IROS 2026 β one 3D flow-matching policy for both single- and dual-arm manipulation, ~30Γ faster than 3D-diffusion policies, +41.4% over the prior best on bimanual PerAct2. Companions: IROS 2026 survey Β· VLA Architectures Β· Real-Time Execution Β· World Models.

3DFA combines flow matching for trajectory prediction with 3D pretrained visual scene representations for learning from demonstration, using 3D relative attention between action and visual tokens during denoising (building on 3D diffusion single-arm policies). The headline is a unified architecture that handles single and dual-arm without separate designs: a frozen image encoder lifts image+depth into 3D scene tokens, and left/right noised trajectory tokens are denoised jointly by a Transformer that attends over scene, proprioception, and language tokens.
- ~30Γ faster training and inference than prior 3D-diffusion policies, via flow matching + system-level and architectural optimizations β without sacrificing performance. Concretely on PerAct2: training 21 days β 16 hours, inference 0.5 Hz β 18.2 Hz (the latency answer for 3D policies, cf. Real-Time Execution).
- 3D relative attention between action tokens and 3D visual tokens β the inductive bias that grounds actions in scene geometry.
- Dense end-effector trajectory prediction in the unimanual case β eliminates motion planning.
- Bimanual PerAct2: new state of the art, beating the next-best by an absolute +41.4%.
- Real-world: surpasses baselines with up to 1000Γ more parameters and far more pretraining.
- Unimanual: new SOTA on 74 RLBench tasks by directly predicting dense EE trajectories.
- Ablations confirm the design choices drive both effectiveness and efficiency.
3DFA is IROS 2026's strongest evidence for the survey Β§5.3 insight that the bimanual frontier is coordination architecture, not just adding a second arm: a single 3D policy unifies single/dual-arm, and the win comes from 3D-geometry grounding + flow-matching efficiency, not scale (it beats models 1000Γ larger). It also lands squarely in the geometry-grounded manipulation thread (3D-foundation-aligned VLA) and the flow-matching efficiency thread β a small, fast, geometry-aware policy beating big ones is the counter-narrative to pure scaling.
Limitations (reviewer): relies on a 3D visual representation / calibrated 3D input (not raw RGB VLA); PerAct2/RLBench are the eval substrate; no language-instruction generalization tested (it's a demonstration policy, not a VLA with a VLM backbone).
- Official program: IROS 2026 (paper #4030) Β· survey: IROS 2026
- Related: EquiBim (the other bimanual page) Β· VLA Architectures Β· Real-Time Execution Β· Humanoid VLA
β Back to IROS 2026 survey Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)