Skip to content

ICLR 2026 HAMLET

Heungwoo edited this page Jun 1, 2026 · 3 revisions

HAMLET β€” Switch Your VLA into a History-Aware Policy

Venue: ICLR 2026 (under review) Β· arXiv 2510.00695 Authors: Myungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee, Changyeon Kim, Younggyo Seo, Jinwoo Shin (KAIST Β· UC Berkeley Β· RLWRLD) Category: VLA Architecture β€” Memory Trend tag: Memory / long horizon

Approach diagram

flowchart LR
  Obs[Per-step observation] --> VLM[Pretrained VLM backbone<br/>FROZEN]
  VLM --> MT[Moment token<br/>TCL-initialized event summary]
  MT --> Mem[Memory module<br/>causal-attention Transformer]
  Past[Past moment tokens] --> Mem
  Mem --> Cond[History condition]
  VLM --> Cond
  Cond --> Act[Action head] --> A[Action]
Loading

Problem

Most VLAs are trained on tasks short enough to fit the backbone's context window. Long-horizon tasks (fold this laundry pile; assemble this kit) require memory of events from many seconds ago. Naive trajectory concatenation either explodes context or causes the model to memorize specific trajectories.

Method

Plug-in memory framework on top of a pretrained VLA. Two components:

  • Moment tokens β€” learnable tokens appended to the VLM input that compress the per-step vision-language representation into a compact event summary. They are initialized with time-contrastive learning (TCL): augmented views of a frame act as positives, distant timesteps as hard negatives, so the tokens capture temporally distinctive aspects and filter redundant static content.
  • Memory module β€” a lightweight Transformer that aggregates past moment tokens via causal self-attention into a temporally-informed condition concatenated with the current VLM representation before action prediction.

The VLM backbone stays frozen; only the moment tokens and the memory module are fine-tuned (a parameter-efficient adapter, not a full retrain β€” and not zero-training).

Results

  • Real-world history-dependent tasks (on GR00T N1.5): 76.4% average success, +47.2 pp over the baseline β€” the headline gain, since these tasks genuinely require memory.
  • RoboCasa Kitchen (100-demo): 64.1% β†’ 66.4%.
  • LIBERO: 95.6% β†’ 97.7%.
  • SimplerEnv-Bridge (on CogACT): ~52.1% β†’ ~63.5%, showing the framework generalizes across backbones.

Backbones tested: GR00T N1.5/N1, CogACT, Ο€β‚€ / Ο€β‚€-FAST, plus an OpenVLA autoregressive extension. Ablations: removing the memory module causes the largest drop; removing TCL initialization hurts consistently; a Transformer memory beats RNN/LSTM/GRU variants.

Significance

Plug-and-play design makes long-horizon capability a cheap add-on, not a reason to retrain the whole model. Introduces a principled distinction between context (raw window) and memory (compressed past-but-relevant info).

Links

πŸ“– In-depth cross-paper review

For a comparison with all major VLA memory architectures (MEM, MemoryVLA, MemER, SAM2Act+, etc.) β€” categorized into 6 groups with pros/cons: VLA Memory Architectures Review.

Related pages

  • MemoryVLA (perceptual + cognitive memory bank)

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally