-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 HAMLET
Venue: ICLR 2026 (under review) Β· arXiv 2510.00695 Authors: Myungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee, Changyeon Kim, Younggyo Seo, Jinwoo Shin (KAIST Β· UC Berkeley Β· RLWRLD) Category: VLA Architecture β Memory Trend tag: Memory / long horizon
flowchart LR
Obs[Per-step observation] --> VLM[Pretrained VLM backbone<br/>FROZEN]
VLM --> MT[Moment token<br/>TCL-initialized event summary]
MT --> Mem[Memory module<br/>causal-attention Transformer]
Past[Past moment tokens] --> Mem
Mem --> Cond[History condition]
VLM --> Cond
Cond --> Act[Action head] --> A[Action]
Most VLAs are trained on tasks short enough to fit the backbone's context window. Long-horizon tasks (fold this laundry pile; assemble this kit) require memory of events from many seconds ago. Naive trajectory concatenation either explodes context or causes the model to memorize specific trajectories.
Plug-in memory framework on top of a pretrained VLA. Two components:
- Moment tokens β learnable tokens appended to the VLM input that compress the per-step vision-language representation into a compact event summary. They are initialized with time-contrastive learning (TCL): augmented views of a frame act as positives, distant timesteps as hard negatives, so the tokens capture temporally distinctive aspects and filter redundant static content.
- Memory module β a lightweight Transformer that aggregates past moment tokens via causal self-attention into a temporally-informed condition concatenated with the current VLM representation before action prediction.
The VLM backbone stays frozen; only the moment tokens and the memory module are fine-tuned (a parameter-efficient adapter, not a full retrain β and not zero-training).
- Real-world history-dependent tasks (on GR00T N1.5): 76.4% average success, +47.2 pp over the baseline β the headline gain, since these tasks genuinely require memory.
- RoboCasa Kitchen (100-demo): 64.1% β 66.4%.
- LIBERO: 95.6% β 97.7%.
- SimplerEnv-Bridge (on CogACT): ~52.1% β ~63.5%, showing the framework generalizes across backbones.
Backbones tested: GR00T N1.5/N1, CogACT, Οβ / Οβ-FAST, plus an OpenVLA autoregressive extension. Ablations: removing the memory module causes the largest drop; removing TCL initialization hurts consistently; a Transformer memory beats RNN/LSTM/GRU variants.
Plug-and-play design makes long-horizon capability a cheap add-on, not a reason to retrain the whole model. Introduces a principled distinction between context (raw window) and memory (compressed past-but-relevant info).
For a comparison with all major VLA memory architectures (MEM, MemoryVLA, MemER, SAM2Act+, etc.) β categorized into 6 groups with pros/cons: VLA Memory Architectures Review.
- MemoryVLA (perceptual + cognitive memory bank)
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)