-
Notifications
You must be signed in to change notification settings - Fork 0
Changelog
hwoo.han edited this page Aug 11, 2026
·
23 revisions
Reverse-chronological log of page additions and major updates to this wiki. Dates are page-addition dates; month-level where the exact day is approximate. Back to Home.
- 2026-08-11 β DYNA-2 (in-depth) added to Latest Papers β Dyna Robotics' World-Action Model launch (Aug 10, 2026): a WAM on ~1M h human egocentric video with no robot data in pre-training, joint next-frame+next-action, claiming the first human-to-robot scaling law smooth over 1kβ1M h (~50Γ EgoScale) and 87% vs 46% zero-shot pass over its DYNA-1 VLA. Reviewed with an explicit vendor-claim caveat β company announcement, no technical paper/benchmark/weights. Cross-linked from World-Models, Human-Video-Transfer, Reviews.
- 2026-08-11 β New Latest Papers preprint tracker + two in-depth preprint reviews: Ο-0 (arXiv 2608.06375 β whole-body humanoid World Action Model using reconstruction-free latent future prediction + SONIC control; single model does 11 household loco-manipulation tasks at 81.8% SR / 90.3% progress vs 44.5% / 59.6% for Ο-0; ships the 40 h six-modality Ο-HOME dataset; Fig. 2 embedded) and Stellar VLA (arXiv 2511.18085 β continual imitation learning for a fixed ~1B VLA with 1% replay via a Dirichlet-Process self-evolving knowledge space + knowledge-routed diffusion MoE; SOTA CIL on LIBERO and 90.0% Final SR on real dual-arm; Fig. 2 embedded). Cross-linked from Home, Reviews, sidebar, Humanoid-VLA and World-Models reviews.
- 2026-08-10 β DreamZero (in-depth) β NVIDIA's World Action Models are Zero-shot Policies (arXiv 2602.15922): a 14B video-diffusion World Action Model that jointly predicts video+action, beats SOTA VLAs by >2Γ on unseen-env/unseen-object real-robot evals (62.2% vs 27.4% seen-task progress; VLAs ~0% at matched scale), with a 38Γ inference stack (DreamZero-Flash decoupled noise schedule) for 7 Hz closed-loop control and video-only cross-embodiment transfer (+42% from β€20 min; 30-min new-robot adaptation). Fig. 4 architecture embedded. Cross-linked from World Models review, Home, Reviews, sidebar.
- 2026-08-05 β (fix) RSS 2026 survey hit GitHub's wiki render limit (262 wikilinks / 40 KB) β the full session tables (116 linked rows) moved to the new RSS-2026-Papers index page; the survey keeps a pointer and now renders at ~147 links / 21 KB.
- 2026-08-05 β RSS 2026 coverage completed: 101 new per-paper reference pages generated for every remaining in-scope paper (verbatim program abstract Β· session/authors/program-page metadata Β· related-topic-review links), bringing all 116 in-scope papers to full page coverage. RSS 2026 survey updated: all session tables fully linked (103 rows), the condensed Humanoids/WM/RL/Datasets block expanded into five per-session tables with abstract-first-line glosses (incl. a new hands/tactile picks table), 100 bolded mentions in the themed lists linked, and each of the 8 themed reading lists now opens with a π§ technology-level insight paragraph (reward-source differentiation in RL, the three-camp mechanism choice for human data, prediction-space selection for world models, touch-as-predicted-state, the hand canonicalization playbook, decomposition-beats-end-to-end for humanoids, ordered-tokens vs native-continuation levers, evaluation-as-systems-discipline). RSS hub updated.
- 2026-08-05 β Deep-dive surveys updated with ICML 2026 (previously RSS-only): 14 sections gain ICML evidence from the 99-paper index β architecture (VLANeXt recipe Β· From-Pixels-to-Tokens Oral Β· XR-1 Oral), wiring (MoT dual-systems HALO/LaSTβ Β· Move-Then-Operate Β· LangForce shortcut counter), attention/real-time (the 9-paper efficiency cluster: Reflex 50 Hz streaming Β· GridS β76% FLOPs Β· SpecPrune/EcoVLA Β· XPU two-phase profile Β· latent reasoning β90% latency), RL (VLAC/LAGEA/ReLAM rewards Β· VLA-MBPO/VLAW model-based +39.2% Β· test-time critics), memory (HiMe Β· SOMA Β· CAPS drift), world models (DreamDojo 44k h Β· LAC-WM +46.7% Β· dWorldEval action-token evaluator), dexterous (DexMachina Β· DECO Β· Tabero β70% grip force Β· CTSRL), cross-embodiment (OXE-AugE +24β45% Β· latent motion codes), evaluation (LIBERO-Gen tiers Β· VLA-Arena Β· FixBench Β· TRAP adversarial), human-video (videoβworld-model fourth use). Verdicts revised: forgetting is milder than assumed (VLA-Forgetting Oral) but untested under repeated RL; discrete-token verdict scoped to robot-action auxiliaries (latent-action tokens as VLM supervision are effective); the WM-evaluator action-input gap has its first crack (dWorldEval). Home fold-outs synced.
-
2026-08-05 β Decision-map restructure for scannability: every Home fold-out now leads with a prominent π Deep dive β link (previously buried after the limitations text) and compresses π Trend / βοΈ Approaches /
β οΈ Open into a compact 3-row label table; the 12 detail-page survey sections rewritten to a uniform template β## π State of the Field (updated Aug 2026)+ one-line Verdict quote +π Trend/βοΈ Approaches & trade-offs/ (optionalβ Established findings) /β οΈ Limitations & open problemsβ with content carried over and tightened; the three standalone surveys (Review-Human-Video-Transfer Β· Review-VLA-Evaluation Β· Review-Realtime-Execution) got matching (updated Aug 2026) headers with structure legends.
- 2026-07-27 β Decision-map topics upgraded to survey-report depth on their detail pages: a dated "State of the Field β July 2026" section (trend arc Β· approach taxonomy with trade-offs Β· limitations) added to 12 topic reviews (Review-VLA-Architecture Β· Review-VLM-Action-Connection Β· Review-VLA-Attention Β· Review-Independent-Visual-Representation Β· Review-VLA-Training-Frameworks Β· RL Β· Review-VLA-Memory Β· Review-World-Models Β· Review-Dexterous-Manipulation Β· Review-Cross-Embodiment Β· Review-Humanoid-VLA Β· Review-LBM-Cotraining consolidated-verdict table), and three new topic surveys created for previously page-less questions: Review-Human-Video-Transfer (emergence vs decoupling vs synthesis, with a decision guide), Review-VLA-Evaluation (the indictment, approach taxonomy, emerging norms), Review-Realtime-Execution (RTCβLegato arc, approach comparison, latency-reporting gap). Home fold-outs now link the full surveys; Reviews catalog updated.
-
2026-07-27 β Home Research Decision Map rewritten from a link table into an insight map: all 15 topics are now collapsible entries, each carrying π the trend arc through the latest venues (RSS 2026 / ICML / ICRA / ICLR 2026), βοΈ the competing approaches with definitions and trade-offs, and
β οΈ current limitations β e.g. the flow-vs-AR arc with OAT's revival, the cross-attention-vs-concatenation ablation status, the RL-from-experience production milestone, the human-video emergence-vs-decoupling fork, the co-training verdicts, the evaluation-era shift, and the alignment-as-scaling-precondition reframe for cross-embodiment. All claims sourced from wiki-verified numbers. - 2026-07-25 β Home redesign: stat strip added, Research Decision Map reorganized into four themed tables (π Building Β· β‘ Running & improving Β· π Data & evaluation Β· π¦Ύ Embodiment) with a "current answer in one line" column, emoji section headers, compact 3-row reviews table; the Page Format / Maintenance Rule section moved off Home to the new Maintenance page (footer link).
- 2026-07-25 β Navigation restructure: new Reviews page as the full in-depth-review catalog (topic reviews Β· lab programs Β· per-paper long-forms Β· RSS 2026 figure pages); _Sidebar slimmed from ~250 to ~45 lines, keeping only top-level entries (reviews hub + 6 star topics, model lineages, ML hub, one link per venue year, foundational refs) β per-paper links now live on venue pages and in Reviews; Home Β§3/Β§5 merged into a compact reviews section pointing at the catalog.
- 2026-07-25 β Knowledge Graph updated with RSS 2026: venue node added to the overall map (+ previously-missing ICML), Tactile-VLA/Humanoid-VLA nodes and RSS-labeled edges in the reviews graph, lineage extensions (RECAP marked RSS-oral; RTC β Legato branch + Ξ¨β adoption; new TRI-LBM and LIBERO robustness lines), world-model cluster additions (mimic-video, LDA-1B under a new unified-WM branch, Qwen-RobotWorld), and a new fifth view: the RSS 2026 improvement-loop cluster mapping all six threads to their 15 pages.
- 2026-07-25 β Home Research Decision Map refreshed with RSS 2026 findings: four new question rows (policy improvement from experience Β· human-video-vs-robot-data fork Β· co-training data selection Β· credible evaluation) and five rows updated with RSS evidence (Legato/OAT for real-time inference, LDA-1B/mimic-video for world models, ViTacFormer/CGP for dexterity, cross-hand pages + camera-frame EEF for cross-embodiment, Ξ¨β for humanoids). Header stats and start-here pointer (β RSS survey) updated.
- 2026-07-25 β Remaining six RSS 2026 per-paper pages upgraded to figure-illustrated reviews with confirmed arXiv IDs: RSS-2026-OAT (2602.04215, HarvardΓStanford β desiderata Venn + anytime-decoding chart), RSS-2026-Legato (2602.12978, SJTUΓSpirit AI β smoothness-vs-time scatter + hesitation traces), RSS-2026-Contact-Grounded-Policy (2603.05687, PurdueΓMeta β teleop channels + predicted-contact pipeline; CVPR-workshop Outstanding award noted), RSS-2026-DexGrasp-Zero (2603.16806 β paradigm comparison + unseen-hand deployment grid), RSS-2026-One-Hand (2602.16712, UNC β canonical-hand overview), RSS-2026-LIBERO-X (2602.06556, MeituanΓBeihang β L1βL5 pyramid; 2,520 demos/600 tasks/100 scenes added). Sidebar: RSS added to Conferences + full RSS 2026 section + recent in-depth reviews (Qwen series, Ξ¨β) listed.
- 2026-07-25 β RSS 2026 major papers upgraded to figure-illustrated 1-page reviews: key figures extracted from the arXiv originals (all verified) with explanatory captions added to PI-RECAP (RECAP loop overview), Review-Psi0 (G1 pantry teaser), RSS-2026-LDA-1B (data-tier/objective/results teaser), RSS-2026-Human2Robot-Emergence (emergence-vs-diversity chart), RSS-2026-mimic-video (VLA-vs-VAM + 10Γ chart), RSS-2026-LBM-Cotraining-Study (modality-matrix overview), RSS-2026-ViTacFormer (CVAE architecture + SharpaWave hardware details), RSS-2026-PolaRiS (4-part system + correlation scatter), RSS-2026-HoMMI (collection/gap/skills). arXiv IDs confirmed for all nine.
- 2026-07-25 β RSS 2026 β VLA & Manipulation Survey β Sydney, Jul 13β17; 210 accepted papers, ~116 manipulation/hand/humanoid in scope, all abstracts verified against the official program. Six threads: RL-from-experience for VLAs (Ο*0.6/RECAP flagship), human-video transfer (emergence vs decoupling), video/world models vs VLA backbones, contact-as-representation, cross-embodiment dexterous hands, evaluation infrastructure. + RSS venue hub, 14 new per-paper pages (LDA-1B Β· H2R-Emergence Β· mimic-video Β· LBM co-training study Β· ViTacFormer Β· DexGrasp-Zero Β· One-Hand Β· Contact-Grounded Policy Β· PolaRiS Β· LIBERO-X Β· OAT Β· Legato Β· HoMMI), and PI-RECAP updated with the RSS camera-ready results.
- 2026-07-25 β Ξ¨β (in-depth) β open humanoid loco-manipulation foundation model (USC PSI Lab Γ NVIDIA, RSS 2026, arXiv 2603.12263): decoupled human-videoβVLM / robot-dataβMM-DiT recipe; 800 h EgoDex + 30 h robot data beats >10Γ corpora (incl. GR00T N1.6) by >40 pp on 8 real Unitree-G1 tasks; full-paper verified.
- 2026-07-24 β Qwen-VLA (in-depth) (minor update) β re-verified against arXiv v2 (Jun 1, 2026): added the new no-T2A baseline (60.9% β T2A worth +10.2 pp) and the v2 clarification that T2A shares the downstream action representation (chunk-first-frame delta EEF); T2A ablation numbers aligned to v2's one-decimal precision. All benchmark tables unchanged between versions.
- 2026-07-24 β VLM4VLA (in-depth) (minor update) β re-verified against arXiv v2 (May 2026): corrected Ο0 Calvin Task-3 (0.786 β 0.686, the only substantive v1βv2 change); added version history and explicit Tsinghua Γ Qwen affiliation split to the header.
- 2026-07-24 β Qwen-RobotNav (in-depth) β navigation specialist on Qwen3-VL: parameterized observation interface (token budget Β· temporal decay Β· camera weights) with training-time randomization, 15.6M-sample corpus, agentic two-tier system (Qwen3.6-Plus planner + evidence notebook) with EQA SOTA; beats Qwen-VLA's nav numbers by ~15 pp.
- 2026-07-24 β Qwen-RobotWorld (in-depth) β language-actioned video world model: 20B double-stream MMDiT + frozen Qwen2.5-VL action encoder, 8.6M-pair EWK corpus with five-layer action-language annotation, Scene2Robot H2R editing; 1st on EWMBench/DreamGen Bench. Program page updated with both.
- 2026-07-24 β Qwen Team's VLA Program β cross-paper review of the Qwen team's VLA line (VLM4VLA β Qwen-VLA β Qwen-Robot Suite incl. RobotNav/RobotWorld): the shared doctrine (Qwen3.5-4B Β· flow matching Β· Ξ»=0.1 VL co-training Β· synthetic-data scaling Β· language-as-interface), the diagnostic-to-flagship trace, and the internal contradictions between the two flagship VLAs.
- 2026-07-24 β Qwen-RobotManip (in-depth) β the Qwen team's second VLA (arXiv 2606.17846): alignment-first scaling thesis (80-dim canonical rep Β· camera-frame delta EEF + CaPE Β· in-context adaptation), ~38,100 h open-data-only corpus (24,808 h human-to-robot synthesis across 15 platforms), new RoboTwin-IF / RoboTwin-XE OOD benchmarks, RoboChallenge Table30-v1 generalist #1.
- 2026-06-11 β Action Space: EEF vs Joint (in-depth) β taxonomy (joint vs EEF Β· absolute vs delta Β· chunking), the EEF-vs-joint verdict, and the O(k)-vs-O(1) chunk-wise-delta result.
- 2026-06-11 β VLA Training Frameworks β cross-paper review of StarVLA (Lego-modular) vs TRI VLA Foundry (LLMβVLMβVLA); structure, supported features, pros/cons.
- 2026-06-10 β RoboMME (in-depth) β memory-implementation breakdown (3 representations Γ 3 integration mechanisms on Ο0.5), full results tables, and the FrameSamp+Modulator code analysis.
- 2026-06-08 β ICML 2026 index β 99 manipulation papers in 10 categories (abstract-verified), the π Top-20 synthesized ranking, and 16 new per-paper pages.
- 2026-06 β Qwen-VLA β the Qwen team's first dedicated VLA.
- 2026-06 β DuoCore-FS β Astribot's parallel fast-slow whole-body stack.
- 2026-06 β Tactile VLA β touch/force-grounded VLAs: architecture Γ sensor hardware Γ 2026 trends.
- 2026-06 β World Models β model-centric robotics world-model taxonomy (incl. WM + inverse-dynamics).
- 2026-06 β WAM vs VLA Robustness β first controlled world-model-vs-VLA benchmark.
- ICRA 2026 Survey β Vienna; record 5,088 submissions; 728 manip/VLA papers; deployment/sensor-centric (+12 topic analyses).
- CVPR 2026 Survey β Denver; VLA is now the modal manipulation contribution (+29 per-paper pages).
- Genesis GENE-26.5 Β· AsyncVLA Β· LBM Co-training (TRI) Β· Steerable Policies Β· GR00T N1βN1.7.
- OmniVTA (visuo-tactile WM) β visuo-tactile world model.
- Goal-Image Conditioning Β· System 0/1/2 Β· VLA Attention Β· ML foundations.
- Per-paper / per-series deep-dives: Ο0.7 Β· Ο0.6 Β· Ο Series Evolution Β· Fast-in-Slow Β· Discrete Diffusion VLA Β· VLM4VLA.
- Core topic reviews: VLA Architectures Β· VLMβAction Connection Β· VLA Memory Β· Dexterous Manipulation Β· Cross-Embodiment Β· RL for VLA.
- Venue surveys: NeurIPS 2025 Β· CoRL 2025 Β· IROS 2025 Β· ICLR 2026 Β· CVPR 2025.
- Foundational references: OpenVLA Β· ReKep Β· AgiBot World Colosseo Β· RoboBrain 2.0.
How to maintain: when adding a page, prepend a dated bullet under the current month. Use exact YYYY-MM-DD when known, else YYYY-MM. Keep Home Β§2 to the ~6 most recent; everything else lives here.
β Back to Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)