Skip to content

History / Review RoboTTT

Revisions

  • Add in-depth review of NVIDIA RoboTTT (context scaling via TTT in GR00T N1.7) Review-RoboTTT (arXiv 2607.15275, NVIDIA GEAR + Stanford + UT Austin): 8K-timestep visuomotor context at constant latency by adding Test-Time- Training fast-weight layers to GR00T N1.7's DiT action head. Detailed GR00T implementation section: TTT layer after self/cross-attn in each of 16 DiT layers (~10M each -> ~690M), tanh-gated to preserve pretrained skills; register tokens (N=16) carry compressed VL history through TTT while VL tokens bypass; fast weights = 2-layer GeLU MLP updated per step (W_t <- W_{t-1} - eta*grad MSE(f(K),V)), read via Q; training recipe = flow matching + sequence action forcing (per-step tau) + TBPTT (fast weights carried, gradients detached at segment boundaries); 30 Hz on RTX 5090 (YAM bimanual). Results, new capabilities (one-shot in-context video imitation, DAgger- distillation self-improvement, perturbation robustness), limitations. Cross-linked from Reviews, Home lab programs, and Review-GR00T-Series. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @Heungwoo Heungwoo committed Aug 19, 2026