Skip to content

Changelog

hwoo.han edited this page Aug 11, 2026 · 23 revisions

Changelog

Reverse-chronological log of page additions and major updates to this wiki. Dates are page-addition dates; month-level where the exact day is approximate. Back to Home.

2026-08

  • 2026-08-11 β€” DYNA-2 (in-depth) added to Latest Papers β€” Dyna Robotics' World-Action Model launch (Aug 10, 2026): a WAM on ~1M h human egocentric video with no robot data in pre-training, joint next-frame+next-action, claiming the first human-to-robot scaling law smooth over 1kβ†’1M h (~50Γ— EgoScale) and 87% vs 46% zero-shot pass over its DYNA-1 VLA. Reviewed with an explicit vendor-claim caveat β€” company announcement, no technical paper/benchmark/weights. Cross-linked from World-Models, Human-Video-Transfer, Reviews.
  • 2026-08-11 β€” New Latest Papers preprint tracker + two in-depth preprint reviews: Ο‰-0 (arXiv 2608.06375 β€” whole-body humanoid World Action Model using reconstruction-free latent future prediction + SONIC control; single model does 11 household loco-manipulation tasks at 81.8% SR / 90.3% progress vs 44.5% / 59.6% for ψ-0; ships the 40 h six-modality Ο‰-HOME dataset; Fig. 2 embedded) and Stellar VLA (arXiv 2511.18085 β€” continual imitation learning for a fixed ~1B VLA with 1% replay via a Dirichlet-Process self-evolving knowledge space + knowledge-routed diffusion MoE; SOTA CIL on LIBERO and 90.0% Final SR on real dual-arm; Fig. 2 embedded). Cross-linked from Home, Reviews, sidebar, Humanoid-VLA and World-Models reviews.
  • 2026-08-10 β€” DreamZero (in-depth) β€” NVIDIA's World Action Models are Zero-shot Policies (arXiv 2602.15922): a 14B video-diffusion World Action Model that jointly predicts video+action, beats SOTA VLAs by >2Γ— on unseen-env/unseen-object real-robot evals (62.2% vs 27.4% seen-task progress; VLAs ~0% at matched scale), with a 38Γ— inference stack (DreamZero-Flash decoupled noise schedule) for 7 Hz closed-loop control and video-only cross-embodiment transfer (+42% from ≀20 min; 30-min new-robot adaptation). Fig. 4 architecture embedded. Cross-linked from World Models review, Home, Reviews, sidebar.
  • 2026-08-05 β€” (fix) RSS 2026 survey hit GitHub's wiki render limit (262 wikilinks / 40 KB) β€” the full session tables (116 linked rows) moved to the new RSS-2026-Papers index page; the survey keeps a pointer and now renders at ~147 links / 21 KB.
  • 2026-08-05 β€” RSS 2026 coverage completed: 101 new per-paper reference pages generated for every remaining in-scope paper (verbatim program abstract Β· session/authors/program-page metadata Β· related-topic-review links), bringing all 116 in-scope papers to full page coverage. RSS 2026 survey updated: all session tables fully linked (103 rows), the condensed Humanoids/WM/RL/Datasets block expanded into five per-session tables with abstract-first-line glosses (incl. a new hands/tactile picks table), 100 bolded mentions in the themed lists linked, and each of the 8 themed reading lists now opens with a 🧠 technology-level insight paragraph (reward-source differentiation in RL, the three-camp mechanism choice for human data, prediction-space selection for world models, touch-as-predicted-state, the hand canonicalization playbook, decomposition-beats-end-to-end for humanoids, ordered-tokens vs native-continuation levers, evaluation-as-systems-discipline). RSS hub updated.
  • 2026-08-05 β€” Deep-dive surveys updated with ICML 2026 (previously RSS-only): 14 sections gain ICML evidence from the 99-paper index β€” architecture (VLANeXt recipe Β· From-Pixels-to-Tokens Oral Β· XR-1 Oral), wiring (MoT dual-systems HALO/LaSTβ‚€ Β· Move-Then-Operate Β· LangForce shortcut counter), attention/real-time (the 9-paper efficiency cluster: Reflex 50 Hz streaming Β· GridS βˆ’76% FLOPs Β· SpecPrune/EcoVLA Β· XPU two-phase profile Β· latent reasoning βˆ’90% latency), RL (VLAC/LAGEA/ReLAM rewards Β· VLA-MBPO/VLAW model-based +39.2% Β· test-time critics), memory (HiMe Β· SOMA Β· CAPS drift), world models (DreamDojo 44k h Β· LAC-WM +46.7% Β· dWorldEval action-token evaluator), dexterous (DexMachina Β· DECO Β· Tabero βˆ’70% grip force Β· CTSRL), cross-embodiment (OXE-AugE +24–45% Β· latent motion codes), evaluation (LIBERO-Gen tiers Β· VLA-Arena Β· FixBench Β· TRAP adversarial), human-video (videoβ†’world-model fourth use). Verdicts revised: forgetting is milder than assumed (VLA-Forgetting Oral) but untested under repeated RL; discrete-token verdict scoped to robot-action auxiliaries (latent-action tokens as VLM supervision are effective); the WM-evaluator action-input gap has its first crack (dWorldEval). Home fold-outs synced.
  • 2026-08-05 β€” Decision-map restructure for scannability: every Home fold-out now leads with a prominent πŸ“„ Deep dive β†’ link (previously buried after the limitations text) and compresses πŸ“ˆ Trend / βš–οΈ Approaches / ⚠️ Open into a compact 3-row label table; the 12 detail-page survey sections rewritten to a uniform template β€” ## πŸ—“ State of the Field (updated Aug 2026) + one-line Verdict quote + πŸ“ˆ Trend / βš–οΈ Approaches & trade-offs / (optional βœ… Established findings) / ⚠️ Limitations & open problems β€” with content carried over and tightened; the three standalone surveys (Review-Human-Video-Transfer Β· Review-VLA-Evaluation Β· Review-Realtime-Execution) got matching (updated Aug 2026) headers with structure legends.

2026-07

  • 2026-07-27 β€” Decision-map topics upgraded to survey-report depth on their detail pages: a dated "State of the Field β€” July 2026" section (trend arc Β· approach taxonomy with trade-offs Β· limitations) added to 12 topic reviews (Review-VLA-Architecture Β· Review-VLM-Action-Connection Β· Review-VLA-Attention Β· Review-Independent-Visual-Representation Β· Review-VLA-Training-Frameworks Β· RL Β· Review-VLA-Memory Β· Review-World-Models Β· Review-Dexterous-Manipulation Β· Review-Cross-Embodiment Β· Review-Humanoid-VLA Β· Review-LBM-Cotraining consolidated-verdict table), and three new topic surveys created for previously page-less questions: Review-Human-Video-Transfer (emergence vs decoupling vs synthesis, with a decision guide), Review-VLA-Evaluation (the indictment, approach taxonomy, emerging norms), Review-Realtime-Execution (RTCβ†’Legato arc, approach comparison, latency-reporting gap). Home fold-outs now link the full surveys; Reviews catalog updated.
  • 2026-07-27 β€” Home Research Decision Map rewritten from a link table into an insight map: all 15 topics are now collapsible entries, each carrying πŸ“ˆ the trend arc through the latest venues (RSS 2026 / ICML / ICRA / ICLR 2026), βš–οΈ the competing approaches with definitions and trade-offs, and ⚠️ current limitations β€” e.g. the flow-vs-AR arc with OAT's revival, the cross-attention-vs-concatenation ablation status, the RL-from-experience production milestone, the human-video emergence-vs-decoupling fork, the co-training verdicts, the evaluation-era shift, and the alignment-as-scaling-precondition reframe for cross-embodiment. All claims sourced from wiki-verified numbers.
  • 2026-07-25 β€” Home redesign: stat strip added, Research Decision Map reorganized into four themed tables (πŸ— Building Β· ⚑ Running & improving Β· πŸ“Š Data & evaluation Β· 🦾 Embodiment) with a "current answer in one line" column, emoji section headers, compact 3-row reviews table; the Page Format / Maintenance Rule section moved off Home to the new Maintenance page (footer link).
  • 2026-07-25 β€” Navigation restructure: new Reviews page as the full in-depth-review catalog (topic reviews Β· lab programs Β· per-paper long-forms Β· RSS 2026 figure pages); _Sidebar slimmed from ~250 to ~45 lines, keeping only top-level entries (reviews hub + 6 star topics, model lineages, ML hub, one link per venue year, foundational refs) β€” per-paper links now live on venue pages and in Reviews; Home Β§3/Β§5 merged into a compact reviews section pointing at the catalog.
  • 2026-07-25 β€” Knowledge Graph updated with RSS 2026: venue node added to the overall map (+ previously-missing ICML), Tactile-VLA/Humanoid-VLA nodes and RSS-labeled edges in the reviews graph, lineage extensions (RECAP marked RSS-oral; RTC β†’ Legato branch + Ξ¨β‚€ adoption; new TRI-LBM and LIBERO robustness lines), world-model cluster additions (mimic-video, LDA-1B under a new unified-WM branch, Qwen-RobotWorld), and a new fifth view: the RSS 2026 improvement-loop cluster mapping all six threads to their 15 pages.
  • 2026-07-25 β€” Home Research Decision Map refreshed with RSS 2026 findings: four new question rows (policy improvement from experience Β· human-video-vs-robot-data fork Β· co-training data selection Β· credible evaluation) and five rows updated with RSS evidence (Legato/OAT for real-time inference, LDA-1B/mimic-video for world models, ViTacFormer/CGP for dexterity, cross-hand pages + camera-frame EEF for cross-embodiment, Ξ¨β‚€ for humanoids). Header stats and start-here pointer (β†’ RSS survey) updated.
  • 2026-07-25 β€” Remaining six RSS 2026 per-paper pages upgraded to figure-illustrated reviews with confirmed arXiv IDs: RSS-2026-OAT (2602.04215, HarvardΓ—Stanford β€” desiderata Venn + anytime-decoding chart), RSS-2026-Legato (2602.12978, SJTUΓ—Spirit AI β€” smoothness-vs-time scatter + hesitation traces), RSS-2026-Contact-Grounded-Policy (2603.05687, PurdueΓ—Meta β€” teleop channels + predicted-contact pipeline; CVPR-workshop Outstanding award noted), RSS-2026-DexGrasp-Zero (2603.16806 β€” paradigm comparison + unseen-hand deployment grid), RSS-2026-One-Hand (2602.16712, UNC β€” canonical-hand overview), RSS-2026-LIBERO-X (2602.06556, MeituanΓ—Beihang β€” L1–L5 pyramid; 2,520 demos/600 tasks/100 scenes added). Sidebar: RSS added to Conferences + full RSS 2026 section + recent in-depth reviews (Qwen series, Ξ¨β‚€) listed.
  • 2026-07-25 β€” RSS 2026 major papers upgraded to figure-illustrated 1-page reviews: key figures extracted from the arXiv originals (all verified) with explanatory captions added to PI-RECAP (RECAP loop overview), Review-Psi0 (G1 pantry teaser), RSS-2026-LDA-1B (data-tier/objective/results teaser), RSS-2026-Human2Robot-Emergence (emergence-vs-diversity chart), RSS-2026-mimic-video (VLA-vs-VAM + 10Γ— chart), RSS-2026-LBM-Cotraining-Study (modality-matrix overview), RSS-2026-ViTacFormer (CVAE architecture + SharpaWave hardware details), RSS-2026-PolaRiS (4-part system + correlation scatter), RSS-2026-HoMMI (collection/gap/skills). arXiv IDs confirmed for all nine.
  • 2026-07-25 β€” RSS 2026 β€” VLA & Manipulation Survey β€” Sydney, Jul 13–17; 210 accepted papers, ~116 manipulation/hand/humanoid in scope, all abstracts verified against the official program. Six threads: RL-from-experience for VLAs (Ο€*0.6/RECAP flagship), human-video transfer (emergence vs decoupling), video/world models vs VLA backbones, contact-as-representation, cross-embodiment dexterous hands, evaluation infrastructure. + RSS venue hub, 14 new per-paper pages (LDA-1B Β· H2R-Emergence Β· mimic-video Β· LBM co-training study Β· ViTacFormer Β· DexGrasp-Zero Β· One-Hand Β· Contact-Grounded Policy Β· PolaRiS Β· LIBERO-X Β· OAT Β· Legato Β· HoMMI), and PI-RECAP updated with the RSS camera-ready results.
  • 2026-07-25 β€” Ξ¨β‚€ (in-depth) β€” open humanoid loco-manipulation foundation model (USC PSI Lab Γ— NVIDIA, RSS 2026, arXiv 2603.12263): decoupled human-videoβ†’VLM / robot-dataβ†’MM-DiT recipe; 800 h EgoDex + 30 h robot data beats >10Γ— corpora (incl. GR00T N1.6) by >40 pp on 8 real Unitree-G1 tasks; full-paper verified.
  • 2026-07-24 β€” Qwen-VLA (in-depth) (minor update) β€” re-verified against arXiv v2 (Jun 1, 2026): added the new no-T2A baseline (60.9% β†’ T2A worth +10.2 pp) and the v2 clarification that T2A shares the downstream action representation (chunk-first-frame delta EEF); T2A ablation numbers aligned to v2's one-decimal precision. All benchmark tables unchanged between versions.
  • 2026-07-24 β€” VLM4VLA (in-depth) (minor update) β€” re-verified against arXiv v2 (May 2026): corrected Ο€0 Calvin Task-3 (0.786 β†’ 0.686, the only substantive v1β†’v2 change); added version history and explicit Tsinghua Γ— Qwen affiliation split to the header.
  • 2026-07-24 β€” Qwen-RobotNav (in-depth) β€” navigation specialist on Qwen3-VL: parameterized observation interface (token budget Β· temporal decay Β· camera weights) with training-time randomization, 15.6M-sample corpus, agentic two-tier system (Qwen3.6-Plus planner + evidence notebook) with EQA SOTA; beats Qwen-VLA's nav numbers by ~15 pp.
  • 2026-07-24 β€” Qwen-RobotWorld (in-depth) β€” language-actioned video world model: 20B double-stream MMDiT + frozen Qwen2.5-VL action encoder, 8.6M-pair EWK corpus with five-layer action-language annotation, Scene2Robot H2R editing; 1st on EWMBench/DreamGen Bench. Program page updated with both.
  • 2026-07-24 β€” Qwen Team's VLA Program β€” cross-paper review of the Qwen team's VLA line (VLM4VLA β†’ Qwen-VLA β†’ Qwen-Robot Suite incl. RobotNav/RobotWorld): the shared doctrine (Qwen3.5-4B Β· flow matching Β· Ξ»=0.1 VL co-training Β· synthetic-data scaling Β· language-as-interface), the diagnostic-to-flagship trace, and the internal contradictions between the two flagship VLAs.
  • 2026-07-24 β€” Qwen-RobotManip (in-depth) β€” the Qwen team's second VLA (arXiv 2606.17846): alignment-first scaling thesis (80-dim canonical rep Β· camera-frame delta EEF + CaPE Β· in-context adaptation), ~38,100 h open-data-only corpus (24,808 h human-to-robot synthesis across 15 platforms), new RoboTwin-IF / RoboTwin-XE OOD benchmarks, RoboChallenge Table30-v1 generalist #1.

2026-06

  • 2026-06-11 β€” Action Space: EEF vs Joint (in-depth) β€” taxonomy (joint vs EEF Β· absolute vs delta Β· chunking), the EEF-vs-joint verdict, and the O(k)-vs-O(1) chunk-wise-delta result.
  • 2026-06-11 β€” VLA Training Frameworks β€” cross-paper review of StarVLA (Lego-modular) vs TRI VLA Foundry (LLMβ†’VLMβ†’VLA); structure, supported features, pros/cons.
  • 2026-06-10 β€” RoboMME (in-depth) β€” memory-implementation breakdown (3 representations Γ— 3 integration mechanisms on Ο€0.5), full results tables, and the FrameSamp+Modulator code analysis.
  • 2026-06-08 β€” ICML 2026 index β€” 99 manipulation papers in 10 categories (abstract-verified), the πŸ… Top-20 synthesized ranking, and 16 new per-paper pages.
  • 2026-06 β€” Qwen-VLA β€” the Qwen team's first dedicated VLA.
  • 2026-06 β€” DuoCore-FS β€” Astribot's parallel fast-slow whole-body stack.
  • 2026-06 β€” Tactile VLA β€” touch/force-grounded VLAs: architecture Γ— sensor hardware Γ— 2026 trends.
  • 2026-06 β€” World Models β€” model-centric robotics world-model taxonomy (incl. WM + inverse-dynamics).
  • 2026-06 β€” WAM vs VLA Robustness β€” first controlled world-model-vs-VLA benchmark.

2026-05

Earlier (2026-Q1 and before)


How to maintain: when adding a page, prepend a dated bullet under the current month. Use exact YYYY-MM-DD when known, else YYYY-MM. Keep Home Β§2 to the ~6 most recent; everything else lives here.

← Back to Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally