-
Notifications
You must be signed in to change notification settings - Fork 0
ICLR 2026 MemoryVLA
Venue: ICLR 2026 Category: VLA Architecture β Memory Trend tag: Memory / long horizon
flowchart LR
Obs[Observation] --> VLM[Prismatic-7B VLM<br/>DINOv2+SigLIP + LLaMA-7B]
VLM --> P[Perceptual tokens<br/>256 compressed visual tokens]
VLM --> C[Cognitive token<br/>EOS-position latent embedding]
P --> Bank[Perceptual-Cognitive Memory Bank]
C --> Bank
Bank --> Q[Retrieve + gated fusion<br/>at each step]
Q --> Pol[DiT action head<br/>DDIM, 10 steps, ~300M]
Pol --> A[Action chunk]
Like HAMLET: long-horizon manipulation is non-Markovian and needs memory beyond a typical attention window. MemoryVLA distinguishes perceptual memory (low-level visual details) from cognitive memory (high-level semantics), both stored as latent representations consolidated from working memory.
Built on a Prismatic-7B VLM (DINOv2 + SigLIP vision encoders, LLaMA-7B), further pretrained on Open-X Embodiment, with a diffusion Transformer (DiT) action head (DDIM, 10 steps, ~300M params) β a CogACT-style backbone.
Per step the VLM produces working memory: 256 perceptual tokens (SE-bottleneck-compressed visual features) and a single cognitive token taken from the EOS-position output (
A Perceptual-Cognitive Memory Bank stores consolidated entries:
- Retrieval: current tokens query the bank via scaled dot-product attention (with sinusoidal timestep positional encodings on stored entries).
-
Fusion: a learned gate adaptively blends retrieved and current tokens,
$\tilde{x} = g^x \odot H^x + (1-g^x)\odot x$ (MLPβsigmoid gate). - Consolidation: on capacity overflow, adjacent entries with highest cosine similarity are merged by averaging (token-merge), beating FIFO.
On SimplerEnv: 71.9% Bridge (+14.6 vs baseline), 72.7% Fractal, 96.5% LIBERO; 41.2% on Mikasa-Robo (+11.8). Across 12 real-world tasks 84.0% average, with +26 on long-horizon tasks vs the SOTA baseline. Baselines include CogACT, Οβ, OpenVLA, Octo, TraceVLA, RoboVLMs.
Decisive ablations (SimplerEnv-Bridge): combining both banks (71.9%) beats cognitive-only (63.5%) or perceptual-only (64.6%); gated fusion > addition (71.9 vs 67.7); token-merge consolidation > FIFO (71.9 vs 66.7); memory length 16 is the sweet spot.
Mirrors a human-like memory split (perceptual β cognitive) and shows that consolidating both low-level and high-level latent memory into a retrievable bank substantially improves long-horizon, temporally dependent manipulation over context-only and single-bank baselines. (The cognitive memory is a latent embedding, not inspectable text β it is not positioned as a "debuggable/readable" memory.)
- arXiv:2508.19236
- Project page
- ICLR 2026 listing
For a comparison with all major VLA memory architectures (MEM, HAMLET, MemER, SAM2Act+, etc.) β categorized into 6 groups with pros/cons: VLA Memory Architectures Review.
β Back to ICLR-2026
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)