Skip to content

IROS 2026 TempoFit

hwoo.han edited this page Sep 7, 2026 · 2 revisions

IROS 2026 β€” TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon VLA

Venue: IROS 2026 (Pittsburgh) Β· paper #1170 Β· Xi'an Jiaotong-Liverpool University (Sun, Yang, Zhang, Ma, Wu, … Chen). Paper: arXiv 2603.07647 (Mar 2026). The training-free memory retrofit of IROS 2026 β€” give a frozen VLA history by reusing its own prefix-attention K/V as a content-addressable runtime state β€” no new tokens, no trainable modules. Companions: VLA Memory Β· In-Context Imitation Β· RoboTTT Β· IROS 2026 survey.

TempoFit β€” (A) a frozen VLA runs per timestep; at one selected memory layer a compact memory state is carried forward across timesteps tβˆ’2 β†’ tβˆ’1 β†’ t; (B) inside that layer: current K/V are stored to a FIFO history, parameter-free K-to-K retrieval produces history logits that a Frame-Gap Temporal Bias (recency-weighted) adjusts, softmax-pools the history K/V, and injects them via norm-preserving residual loading before self-attention (method figure from Sun et al., arXiv 2603.07647, Β© the authors)

1. Problem

Pretrained VLAs are strong at single-step manipulation but their inference is largely memoryless β€” brittle in non-Markovian long-horizon settings (occlusion, state aliasing, subtle post-action changes). Prior fixes inject history either by stacking frames (scales visual tokens + latency, adds near-duplicate pixels) or by learning extra temporal interfaces (require (re-)training, may break the original single-frame inference graph).

2. Method

TempoFit is a training-free temporal retrofit that upgrades frozen VLAs through state-level memory. Key insight: "prefix-attention K/V already forms a model-native, content-addressable runtime state; reusing them across timesteps introduces history without new tokens or trainable modules."

  • Layer-wise FIFO prefix K/V stored at selected intermediate layers.
  • Parameter-free K-to-K retrieval with Frame-Gap Temporal Bias (FGTB) β€” a fixed recency bias (inspired by NLP positional biases) that keeps decisions present-dominant.
  • Pre-attention residual loading with norm-preserving rescaling injects the retrieved context without distribution shift under frozen weights.

3. Results

  • LIBERO-Long: up to +4.0% average success on strong pretrained backbones, at near-real-time latency.
  • Transfers consistently to CALVIN and real-robot long-horizon tasks.
  • No retraining, no architecture change β€” a pure inference-time upgrade.

4. Why it matters (memory lens)

TempoFit is the "KV-cache is the memory" answer in IROS 2026's memory bifurcation (survey Β§5.2, VLA Memory): where structured methods build explicit stores (GaussMemory's 3D-Gaussian scene, PROMPT's memory trees), TempoFit exploits the model's own attention state β€” the cheapest possible retrofit, and training-free (unlike RoboTTT's fast weights, which need TTT training, or frame-stacking, which needs retraining). It's the pragmatic middle path: parametric/content memory with zero training cost. The FGTB recency-bias echoes the "selective history beats full context" lesson from VLA Memory's RSS-2026 trend.

Limitations (reviewer): +4.0% is a modest lift; FIFO/recency bias may drop genuinely old-but-relevant evidence (no learned retrieval); tuned on LIBERO-Long/CALVIN; frozen-weight residual injection assumes the K/V space is stable across the horizon.

5. Links

← Back to IROS 2026 survey Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally