Skip to content

ICLR 2026 EquAct

Heungwoo edited this page Jun 1, 2026 · 3 revisions

EquAct β€” SE(3)-Equivariant Multi-Task Transformer

Venue: ICLR 2026 / Category: VLA Architecture β€” Equivariance / Trend tag: Spatial / 3D for VLA

Approach diagram

flowchart LR
  PC[Point cloud] --> UNet["SE(3)-equivariant point-cloud U-Net<br/>spherical Fourier features"]
  Lang[Language] --> iFiLM["SE(3)-invariant FiLM layers"]
  iFiLM --> UNet
  UNet --> Act[Action]
Loading

Problem

Multi-task manipulation policies tend to break under novel 3D object poses because their networks have no built-in geometric consistency. The paper aims for theoretical generalization to unseen scene transformations rather than relying on data augmentation. EquAct is an open-loop (keyframe / next-best-pose) policy in the RVT/PerAct lineage β€” it predicts discrete end-effector poses from language + point cloud rather than continuous closed-loop control.

Method

EquAct is an SE(3)-equivariant transformer over point clouds. Two ingredients: (1) an efficient SE(3)-equivariant point-cloud U-Net using spherical Fourier features for policy reasoning, and (2) SE(3)-invariant Feature-wise Linear Modulation (iFiLM) layers for language conditioning, so language influences the policy without breaking equivariance.

Results

State-of-the-art across 18 RLBench tasks under SE(3) and SE(2) scene perturbations and across varying training-data sizes, plus four physical robot tasks. Specific per-task numbers (not stated in abstract).

Significance

Stakes out the "principled equivariance" position in 2026's spatial cluster β€” orthogonal to the VLM-centric papers (Spatial Forcing, FALCON, SP-VLA) that bolt spatial signals onto a 2D backbone. The iFiLM trick (language-conditioning that respects equivariance) is the technically novel piece.

Links

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally