Skip to content

RSS 2026 Act2Goal

hwoo.han edited this page Aug 9, 2026 · 2 revisions

Act2Goal: From World Model To General Goal-conditioned Policy

Venue: RSS 2026 (Sydney, Jul 13–17) Β· Session: World Models & Memory Β· paper #15 Authors: Pengfei Zhou, Liliang Chen, Shengcong Chen, Di Chen, Wenzhi Zhao, Rongjun Jin, Guanghui Ren, Jianlan Luo arXiv: 2512.23541 Β· program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Act2Goal method overview (Figure 1 of arXiv 2512.23541, Β© the authors)

The robot receives a visual goal (left, table with flowers in a vase), imagines a sequence of intermediate visual states toward it via a goal-conditioned world model (thought cloud, "Plan"), and executes the planned actions in the real world ("Action") β€” the core imagine-then-act loop of Act2Goal.

Problem

Visual goals are a compact, unambiguous alternative to language for specifying manipulation tasks, but existing goal-conditioned policies predict actions in a single step without modeling task progress, so they degrade on long-horizon tasks and overfit demonstrated state–action mappings. The authors (Agibot Research) ask how a policy can explicitly reason about the visual dynamics needed to reach a distant goal.

Method

Act2Goal couples a Goal-Conditioned World Model (GCWM) β€” built on the Genie Envisioner architecture with language conditioning removed and a goal image concatenated to the observation β€” with an isomorphic flow-matching action expert (1.6B-parameter Video DiT + 160M action DiT, 28 blocks each, joined by cross-attention). Multi-Scale Temporal Hashing (MSTH) splits the imagined trajectory into dense proximal frames (stride-r sampled up to horizon P, with actions at every timestep) for closed-loop control and logarithmically spaced distal frames that anchor global consistency; only proximal actions are executed. Training is two-stage offline imitation (world-model fine-tuning, then joint flow-matching of video and actions), plus optional reward-free online improvement: HER-style hindsight goal relabeling of self-collected rollouts with LoRA-only updates on the edge device.

Results

On four Robotwin 2.0 simulation tasks, Act2Goal beats DP-GC, Ο€0.5-GC, and HyperGoalNet on all Easy-mode tasks (e.g., Pick Bottles 0.80 vs. 0.13 for Ο€0.5-GC) and 3 of 4 Hard-mode tasks. On an AgiBot Genie-01 robot across three real tasks, it scores ID/OOD 0.93/0.90 (whiteboard word writing), 0.75/0.48 (dessert plating), 0.45/0.30 (plug-in), while DP-GC and HyperGoalNet are near 0. Online autonomous improvement converges in ~3 rounds with up to 8Γ— success-rate gains in simulation; on the real OOD plug-in task success climbs from 0.30 to 0.90, and even failed-only rollouts yield improvement.

Significance

First integration of a world model into goal-conditioned policy learning per the authors, and a practical recipe for reward-free on-robot self-improvement (HER + LoRA) that needs no human labels. Fits the wiki's Review-World-Models thread and the deployment-time-adaptation discussion in RL.

← Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally