RoboMME review: correct implementation details from config + code Verified against perceptual-framesamp-modul.yaml and pi0_config.py: - VLM=PaliGemma (SigLIP+Gemma-2B), action expert=Gemma-300M, memory expert=Gemma-150M - best config: 16-frame streaming horizon, 16 mean-pooled tokens/img, 512 budget, pos-emb on, state-emb OFF, memory_token_dim 1024 (2048 for context) - FIX: modulation is applied in EVERY HistoryBlock on the action-expert stream (i==len(xs)-1 indexes experts, not the last layer), before the FFN - add exact-config table Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
RoboMME review: add implementation-detail section (VLM+AE + FrameSamp/Modulator) Trace the best config (perceptual FrameSamp + Memory-as-Modulator) through the official robomme_policy_learning code (openpi/pi0.5 fork): MoT dual-expert backbone, FrameSamp memory-token branch (512-token budget), and the AdaLN modulation path (action features cross-attend memory -> gamma/beta scale-shift on the action expert only, VLM expert untouched). Maps Appendix Eqs 21-23 to history_gemma.py MemoryAttention/MemoryRMSNorm, adds an end-to-end mermaid. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add in-depth RoboMME review (memory implementation + full results + figures) Long-form Review-RoboMME.md covering the cognitively-motivated task taxonomy (temporal/spatial/object/procedural, 16 tasks/4 suites), the MME-VLA suite's memory implementation (3 representations x instantiations x 3 integration mechanisms = 14 variants on a pi0.5 backbone), and full performance comparison (main results, by-functional-requirement, real-world, human study) — all extracted from arXiv 2603.04639. Embeds the paper's Figure 1/2/3 (attributed). Links from the ICML-2026-RoboMME one-pager and sidebar. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>