Skip to content

ICLR 2026 MemoryVLA

Heungwoo edited this page Jun 1, 2026 · 3 revisions

MemoryVLA β€” Perceptual-Cognitive Memory in VLA

Venue: ICLR 2026 Category: VLA Architecture β€” Memory Trend tag: Memory / long horizon

Approach diagram

flowchart LR
  Obs[Observation] --> VLM[Prismatic-7B VLM<br/>DINOv2+SigLIP + LLaMA-7B]
  VLM --> P[Perceptual tokens<br/>256 compressed visual tokens]
  VLM --> C[Cognitive token<br/>EOS-position latent embedding]
  P --> Bank[Perceptual-Cognitive Memory Bank]
  C --> Bank
  Bank --> Q[Retrieve + gated fusion<br/>at each step]
  Q --> Pol[DiT action head<br/>DDIM, 10 steps, ~300M]
  Pol --> A[Action chunk]
Loading

Problem

Like HAMLET: long-horizon manipulation is non-Markovian and needs memory beyond a typical attention window. MemoryVLA distinguishes perceptual memory (low-level visual details) from cognitive memory (high-level semantics), both stored as latent representations consolidated from working memory.

Method

Built on a Prismatic-7B VLM (DINOv2 + SigLIP vision encoders, LLaMA-7B), further pretrained on Open-X Embodiment, with a diffusion Transformer (DiT) action head (DDIM, 10 steps, ~300M params) β€” a CogACT-style backbone.

Per step the VLM produces working memory: 256 perceptual tokens (SE-bottleneck-compressed visual features) and a single cognitive token taken from the EOS-position output ($c \in \mathbb{R}^{1\times d_c}$). Note: both are latent embeddings, not text/symbolic β€” the page previously claimed the cognitive bank was "text-valued" and human-readable, which the paper contradicts.

A Perceptual-Cognitive Memory Bank stores consolidated entries:

  • Retrieval: current tokens query the bank via scaled dot-product attention (with sinusoidal timestep positional encodings on stored entries).
  • Fusion: a learned gate adaptively blends retrieved and current tokens, $\tilde{x} = g^x \odot H^x + (1-g^x)\odot x$ (MLPβ†’sigmoid gate).
  • Consolidation: on capacity overflow, adjacent entries with highest cosine similarity are merged by averaging (token-merge), beating FIFO.

Results

On SimplerEnv: 71.9% Bridge (+14.6 vs baseline), 72.7% Fractal, 96.5% LIBERO; 41.2% on Mikasa-Robo (+11.8). Across 12 real-world tasks 84.0% average, with +26 on long-horizon tasks vs the SOTA baseline. Baselines include CogACT, Ο€β‚€, OpenVLA, Octo, TraceVLA, RoboVLMs.

Decisive ablations (SimplerEnv-Bridge): combining both banks (71.9%) beats cognitive-only (63.5%) or perceptual-only (64.6%); gated fusion > addition (71.9 vs 67.7); token-merge consolidation > FIFO (71.9 vs 66.7); memory length 16 is the sweet spot.

Significance

Mirrors a human-like memory split (perceptual ↔ cognitive) and shows that consolidating both low-level and high-level latent memory into a retrievable bank substantially improves long-horizon, temporally dependent manipulation over context-only and single-bank baselines. (The cognitive memory is a latent embedding, not inspectable text β€” it is not positioned as a "debuggable/readable" memory.)

Links

πŸ“– In-depth cross-paper review

For a comparison with all major VLA memory architectures (MEM, HAMLET, MemER, SAM2Act+, etc.) β€” categorized into 6 groups with pros/cons: VLA Memory Architectures Review.

Related pages

← Back to ICLR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally