Skip to content

ICML 2026 LaST0

hwoo.han edited this page Jun 11, 2026 · 1 revision

LaSTβ‚€: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Models β€” reasoning before acting in a token-efficient latent space

Venue: ICML 2026 (Poster) Category: Reasoning

Problem

Chain-of-thought (CoT) reasoning can help Vision-Language-Action (VLA) models "think before they act," but expressing that reasoning as natural-language tokens is slow and ill-suited to the physical, spatio-temporal nature of manipulation. A robot's relevant intermediate state is not text β€” it is future visual dynamics, 3D structure, and proprioception. LaSTβ‚€ targets a reasoning representation that captures these physical factors while remaining cheap enough for closed-loop control.

Method

LaSTβ‚€ is a VLA framework that reasons before acting in a token-efficient latent spatio-temporal chain-of-thought space. Instead of emitting language CoT, it reasons over a compact latent that encodes:

  • future visual dynamics (how the scene will evolve),
  • 3D structure of the scene, and
  • proprioception (the robot's own state).

To reconcile the differing timescales of deliberation and control, LaSTβ‚€ adopts a Mixture-of-Transformers dual-system design that separates a low-frequency reasoning expert from a high-frequency action expert. The reasoning system operates over the latent spatio-temporal CoT at a slower cadence, while the action system runs at a higher rate to produce control β€” analogous to "System 2" deliberation feeding a fast "System 1" actuator.

flowchart LR
    Obs[Visual + proprio observation] --> R[Low-frequency reasoning expert]
    R --> CoT[Latent spatio-temporal CoT:<br/>future dynamics Β· 3D Β· proprioception]
    CoT --> A[High-frequency action expert]
    A --> ACT[Action]
    subgraph MoT[Mixture-of-Transformers]
      R
      A
    end
Loading

Results

LaSTβ‚€ reports improvements in mean success rate of 13%, 14% and 14% over prior state-of-the-art VLA methods across its evaluated settings. (Detailed per-benchmark tables are not available, as no public arXiv/HTML source was accessible for this entry at the time of writing; the figures above are from the submission's reported headline.)

Significance

LaSTβ‚€ argues that effective robotic reasoning should be latent and physical, not textual: by compressing the chain-of-thought into a spatio-temporal latent over future dynamics, 3D structure, and proprioception, it keeps deliberation token-efficient enough for control. The Mixture-of-Transformers dual-system split β€” slow reasoning, fast action β€” is a clean architectural answer to the timescale mismatch between thinking and acting, and the consistent double-digit success-rate gains suggest the latent-CoT formulation is a meaningful alternative to language-based CoT for VLAs.

Links

  • ICML 2026: poster (link pending)

← Back to ICML-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally