Skip to content

Review Humanoid VLA

hwoo.han edited this page Sep 5, 2026 · 7 revisions

In-Depth Review β€” VLA for Humanoids (whole-body & bipedal loco-manipulation)

Compiled June 2026 Β· Focus: what makes a Vision-Language-Action model for a humanoid technically different from a tabletop-arm VLA β€” specifically (1) mobility / bipedal balance during manipulation, (2) the high-dimensional whole-body action space, and (3) data collection when you cannot simply teleoperate a balancing biped.

Companion reviews: System 0/1/2 architectures Β· GR00T series Β· Dexterous Manipulation Β· Cross-Embodiment Β· VLA Architectures Β· DuoCore-FS.


1. TL;DR β€” humanoid VLA is a layered problem, not a bigger arm

A tabletop VLA maps (image, instruction) β†’ end-effector action. A humanoid cannot be modeled that way, for three reasons that organize this whole review:

  1. The robot can fall over. Every action is conditioned on keeping a ~19–29-DoF underactuated biped balanced. Almost no system lets the VLA emit balance-critical torques directly β€” balance is delegated to a separate RL whole-body controller (WBC) and the VLA talks to it through a latent or a command.
  2. The action space is 2–5Γ— larger and heterogeneous. Arms + hands + waist + neck + legs, mixing position, force, and locomotion commands at different rates. The field's answers are latent action vocabularies, dual-system VLMβ†’fast-policy splits, and stream-specific tokenization β€” not one monolithic head.
  3. You cannot cheaply teleoperate a walking humanoid. So the data engine shifts off teleop toward human video β†’ reconstruction β†’ retargeting, motion-capture β†’ RL-tracking β†’ distillation, and sim-RL teacher β†’ vision student. Teleop survives mostly for stationary upper-body work (Helix, GR00T's real layer).

The single most important structural fact: "humanoid VLA" today is a stack, cleanly described by the System 0/1/2 framing β€” a slow VLM (System 2, 1–10 Hz), a fast visuomotor policy (System 1, 50–200 Hz), and a reflex/balance layer (System 0, 500 Hz–1 kHz). The genuinely-new research is in how the VLA hands off to the balance layer and where the training data comes from.

Headline caveat β€” read before believing any "whole-body" claim: several flagship "humanoid VLAs" are upper-body only and never walk β€” Figure Helix (35-DoF, upper body, no paper, vendor numbers only) and GR00T's public GR-1 deployment (stationary bimanual). True bipedal loco-manipulation VLAs β€” where the policy commands walking/balance β€” remain rare and mostly simulation-trained: LeVERB, WholeBodyVLA, HumanVLA, Humanoid-VLA.


2. Why a humanoid breaks the tabletop-VLA assumptions

Tabletop-VLA assumption Why a humanoid violates it Field's response
The base is fixed; actions don't threaten stability Underactuated biped; any arm motion shifts CoM and can topple it Delegate balance to an RL WBC (System 0/1); VLA emits latent/command, not torque
≀7–14 DoF, one action modality (EE pose) 19–36 DoF spanning legs/waist/arms/hands/neck; mixes position + force + gait Latent action vocab Β· dual-system split Β· stream-specific tokenization Β· hierarchical value decomposition
Teleop scales data linearly A walking biped is unsafe/awkward to teleoperate; balance data isn't demonstrable Human-video reconstruction · mocap→RL→distill · sim teacher→vision student
One control rate (~10–50 Hz) Reasoning (1–10 Hz) and balance (β‰₯500 Hz) differ by ~100Γ— Asynchronous fast-slow inference; per-rate heads
One embodiment per policy Many humanoid morphologies (G1, H1, GR-1, Atlas, AgiBot, Astribot) Per-embodiment MLP projection to a shared latent; cross-embodiment pyramids

This is why humanoid VLA looks like "a VLM bolted onto a learned whole-body controller" rather than "scale imitation until it works."


3. The landscape β€” three layers and four families

The three-layer stack β€” a slow VLM hands off to a fast policy, which hands off to a balance layer:

flowchart TB
  VLM["System 2 Β· slow VLM Β· 1–10 Hz<br/>latent / subtask / motion code"]
  POL["System 1 Β· fast policy Β· 50–200 Hz<br/>diffusion / flow / AR head"]
  WBC["System 0 Β· reflex & balance Β· 500 Hz–1 kHz<br/>RL whole-body controller"]
  ROB["humanoid joints Β· 19–36 DoF"]
  VLM --> POL --> WBC --> ROB
Loading

The four families β€” which approach fits your constraint:

flowchart TB
  Q{Which family?}
  Q -- learned-latent VLA over an RL WBC --> F1[F1. Latent-vocabulary whole-body VLA<br/>LeVERB Β· WholeBodyVLA]
  Q -- VLM then fast-policy split --> F2[F2. Dual-system manipulation VLA<br/>GR00T Β· Helix Β· DuoCore-FS Β· Galaxea]
  Q -- language to motion / limited vision --> F3[F3. Language-to-motion controllers<br/>Lang-To-Loco Β· BFM-Zero Β· Humanoid-GPT]
  Q -- sim or human-video data engine --> F4[F4. Data & sim-to-real engines<br/>VideoMimic Β· VIRAL Β· EgoScale Β· Open-Sim-to-Real]
Loading
  • F1 β€” Latent-vocabulary whole-body VLA. The VLA emits a learned latent action token; an RL WBC decodes it into balanced whole-body joint targets. The cleanest answer to all three axes at once. (LeVERB, WholeBodyVLA.)
  • F2 β€” Dual-system manipulation VLA. Slow VLM + fast policy, heterogeneous DoF handled by per-embodiment MLPs. Strong on manipulation; balance is usually out of scope (upper-body or mobile-base). (GR00T N1–N1.7, Helix, DuoCore-FS, Galaxea G0, Fast-in-Slow.)
  • F3 β€” Language-to-motion / promptable controllers. Bridge language (sometimes not vision) to whole-body motion via latents; these are often the System 0/1 substrate a full VLA sits on. (Lang-To-Loco/RoboGhost, BFM-Zero, Humanoid-GPT, HWC-Loco, HVD, LIFT, UniFP.)
  • F4 β€” Data & sim-to-real engines. Not policies per se but the data-collection machinery that makes the others trainable. (VideoMimic, VIRAL, Open-Sim-to-Real, Gallant, EgoVLA, EgoScale.)

4. Axis 1 β€” Mobility & bipedal balance (the differentiator)

The defining humanoid problem. Three sub-patterns dominate:

4.1 Delegate balance to an RL whole-body controller (near-universal)

The VLA almost never outputs balance torques. Instead a low-level RL WBC owns stability and the VLA steers it:

  • LeVERB (arXiv 2506.13751, Berkeley): a System-1 RL WBC produces dynamics-feasible walking/turning/sitting from a latent instruction; the System-2 vision-language policy never touches joint torques.
  • WholeBodyVLA (ICLR 2026, 2512.11047, AgiBot X2): a Loco-Manipulation-Oriented (LMO) RL policy for advancing / turning / squatting β€” "manipulation-aware locomotion" that moves the base to serve the manipulation goal. This is the explicit bipedal-loco-manip contribution.
  • HumanVLA (NeurIPS 2024, 2406.19972): a state-based teacher trained with goal-conditioned RL + Adversarial Motion Prior (AMP) learns balanced walk-and-carry, then distilled into a VLA student.
  • BFM-Zero (ICLR 2026, 2511.04131, Unitree G1, 29-DoF): a promptable Forward-Backward foundation controller with disturbance rejection (recovers from kicks/pushes never trained on) β€” the substrate a VLA prompts.

4.2 Make locomotion robust and constraint-aware (the controller research)

These are not VLAs but the balance layer the field depends on:

  • HWC-Loco (ICLR 2026, H1/G1, 19-DoF): 3-stage hierarchy with a ZMP-feasibility constraint and robust optimization over an uncertainty set; survives ≀200 N pushes, 15 cm stairs, 20Β° slopes. Learned ZMP-threshold switching beats fixed thresholds (195 vs 459 switches).
  • UniFP (CoRL 2025 Best Paper, 2505.20829): a single unified policy for position and force in legged loco-manipulation β€” infers external force from proprioception (no F/T sensor), no hand-engineered mode switching; +39.5% on contact-rich tasks (wiping, drawers).
  • Gallant (CVPR 2026, 2511.14625): 3-D voxel-grid terrain perception (overhead + lateral constraints, not just a 2-D heightmap) for navigating pipes/shelves/stairs; near-100% staircase success.

4.3 The frequency stack (why this is hard in real time)

Reasoning, control, and balance run ~100Γ— apart: VLM 1–10 Hz β†’ fast policy 50–200 Hz β†’ balance 500 Hz–1 kHz. Industrial systems make this explicit β€” Figure's reported System-0 is a 1 kHz neural prior trained on >1,000 h of human motion capture that replaced ~109k lines of C++ balance code (per Review-System-0-1-2). The async-inference papers (DuoCore-FS 2.6Γ—, Fast-in-Slow 117.7 Hz) exist precisely to keep the fast layer fed while the slow VLM lags.

Reality check: the most-publicized "humanoid" VLAs sidestep Axis 1 entirely. Helix is upper-body only (no walking/balance claim). DuoCore-FS explicitly omits a System-0 balance layer (quasi-static kiosk). GR00T N1's GR-1 deployment is stationary. Bipedal balance during VLA-driven manipulation is demonstrated by a short list β€” LeVERB, WholeBodyVLA, HumanVLA β€” and mostly in simulation.


5. Axis 2 β€” The high-dimensional whole-body action space

DoF roughly doubles-to-quintuples vs a 7-DoF arm: H1 ~19, G1 ~23–29, Astribot S1 25, Figure 35 (upper only), Dexora 36 (bimanual+hands). Five distinct strategies:

Strategy Mechanism Representative Why it helps
A. Latent action vocabulary VLA emits a learned latent token; RL WBC decodes to whole-body joints LeVERB, WholeBodyVLA Decouples semantics from the 30-DoF control problem; no hand-crafted action primitives
B. Dual-system + per-embodiment projection Slow VLM β†’ shared latent β†’ per-embodiment MLP encoders/decoders β†’ fast head GR00T N1–N1.7, Helix, Galaxea G0, Fast-in-Slow One model spans many morphologies; no explicit upper/lower split needed
C. Stream-specific tokenization Residual-VQ-VAE with parallel codebooks per stream (position / SO(3) / gripper) DuoCore-FS (29-dim β†’ 36 fixed tokens, 3.4Γ— more compact than FAST) Fixed-length tokens make AR decoding of whole-body actions tractable and fast
D. Hierarchical value decomposition Decompose the value function along kinematic structure (not the policy) HVD (WB-50 dataset) Credit assignment in high-DoF offline RL while keeping one unified policy
E. Behavioral latent / FB representation Map 29-DoF control into a low-dim latent z (e.g. R²⁡⁢); plan/track in latent BFM-Zero, Lang-To-Loco (64-D motion latent) Few-shot CEM/MPC and language-conditioning become low-dim problems

Notes that matter:

  • Explicit upper/lower-body decomposition is rare. Most systems use a shared latent and let projection layers sort out morphology. WholeBodyVLA's dual-stream head (separate arm-joint vs locomotion-command outputs) is one of the few explicit splits.
  • Frequency mismatch is part of the action-space problem. Helix runs S1 at 200 Hz (80M params) under a 7–9 Hz 7B VLM; FiS reuses only the last 2 of 32 LLaMA blocks at high rate; DuoCore-FS runs an async 25–30 Hz fast head under a 1–3 Hz slow head.
  • Dexterity rides on top. Bimanual 12-DoF hands (EgoVLA's Inspire hands, Dexora's XHAND) compound the DoF count β€” see Dexterous Manipulation for the hand-specific story; the whole-body papers mostly treat hands as another projected stream.

6. Axis 3 β€” Data collection (the real bottleneck)

You cannot teleoperate a balancing biped at scale, so humanoid VLA data comes from four routes. Retargeting (human morphology β†’ robot URDF under joint limits) is the recurring crux.

6.1 Human video β†’ reconstruction β†’ retargeting

The dominant academic route β€” cheap monocular/egocentric human video lifted to robot actions:

  • VideoMimic (CoRL 2025 Best Student Paper, 2505.03729, G1): everyday RGB video β†’ 4-D reconstruction (SMPL mesh + scene, metric scale via joint Levenberg-Marquardt optimization) β†’ retarget β†’ 4-stage sim (mocap pretrain β†’ scene-conditioned tracking β†’ DAgger distill β†’ under-conditioned RL). Removing the mocap-pretrain stage breaks it.
  • EgoVLA (CVPR 2026, 2507.12440, H1 + 2Γ—12-DoF Inspire): ~500K egocentric image-action pairs (HOI4D/HOT3D/HoloAssist/TACO), action = wrist SE(3) + MANO hand params, deployed via IK + hand retargeting. Pretraining lifts long-horizon success 2.22% β†’ 45.93% vs ACT. Zero-shot (no robot finetune) = 0% β€” human video alone is insufficient.
  • EgoScale (arXiv 2602.16710): 20,854 h of action-labeled egocentric video (monocular SLAM + 21-keypoint hand pose + CasADi/IPOPT retargeting) β€” ~20Γ— prior ego-video efforts; feeds GR00T N1.7 with a log-linear scaling law L_val = 0.024 βˆ’ 0.003Β·ln(D) (RΒ²=0.998).
  • WholeBodyVLA learns a Latent Action Model from action-free egocentric video β€” no demonstration labels β€” then a thin teleop layer grounds it.
  • Humanoid-VLA (arXiv 2502.14795): language-motion pre-alignment β†’ egocentric video-conditioned PEFT β†’ self-supervised pseudo-annotation of unlabeled video.

6.2 Motion capture β†’ RL tracking experts β†’ distillation

For locomotion/whole-body motion (where balance, not semantics, is the target):

  • Humanoid-GPT (CVPR 2026, G1): 2-billion-frame retargeted mocap (AMASS + LAFAN1 + Motion-X++ + MotionMillion + in-house), Harmonic-Motion-Embedding diversity clustering β†’ per-cluster RL experts β†’ DAgger into one causal GPT; zero-shot motion tracking at <1.5 ms (TensorRT).
  • BFM-Zero: LAFAN1 (40 motions retargeted to G1), EMD-prioritized sampling, 192 M env-steps off-policy.
  • Lang-To-Loco (RoboGhost): MotionMillion (50,378 seqs β†’ 3,261 stable after filtering); language β†’ 64-D motion latent β†’ diffusion student; retargeting-free.

6.3 Sim-RL teacher β†’ vision student (the sim-to-real engine)

Privileged-state RL in sim, distilled to an RGB policy that transfers zero-shot:

  • VIRAL (CVPR 2026, 2511.15200, NVIDIA/CMU/Berkeley, G1): privileged teacher β†’ RGB student via large-scale tiled-rendering DAgger; compute is decisive (up to 64 GPUs; low-compute runs fail); 54 zero-shot real cycles.
  • Open-Sim-to-Real (CVPR 2026, 2512.01061): staged-reset PPO teacher β†’ RGB student β†’ GRPO finetune for partial observability; humanoid door-opening 31.7% faster than human teleoperators.
  • LeVERB and HumanVLA are entirely synthetic (rendered kinematic demos + RL); no teleop, no real video.

6.4 Teleoperation (the industrial route, survives for stationary/upper-body)

  • Figure Helix: ~500 h multi-robot multi-operator teleop (vendor figure; no paper).
  • GR00T: the "data pyramid" β€” 88 h GR-1 teleop (VIVE + Xsens) at N1, scaling to thousands of hours of real teleop by N1.6 (YAM/AGIBot/G1). The version-over-version story is which layer gets scaled: N1β†’N1.5 synthetic (DreamGen) Β· N1.5β†’N1.6 real-teleop Β· N1.6β†’N1.7 human-video (EgoScale).
  • Dexora: hybrid exoskeleton backpack + Apple Vision Pro teleop (gross arm + markerless fingers); 100 K sim + 10 K real episodes.
  • DuoCore-FS: just 10.22 h teleop, single kiosk task β€” efficient but narrow.

The data takeaway: academic whole-body VLAs lean synthetic / human-video; industrial manipulation VLAs lean teleop; GR00T is the only line that fuses all three at scale. Retargeting quality (metric-scale reconstruction, IK under joint limits) is the silent determinant of whether human-video data transfers.


7. Per-venue trends (2025–2026)

  • CoRL 2025 β€” human-video-as-cross-embodiment-source goes mainstream. VideoMimic (Best Student Paper) and UniFP (Best Paper) frame the year: monocular human video β†’ balanced humanoid skills, and unified position-force loco-manipulation.
  • NeurIPS 2025 β€” async dual-system inference. Fast-in-Slow embeds the fast policy inside the VLM's last blocks (117.7 Hz) β€” the architectural enabler for running a slow reasoner over a fast humanoid.
  • ICLR 2026 β€” the whole-body controller substrate matures. A dense cluster of non-VLA controllers a VLA sits on: BFM-Zero (promptable FB foundation model), WholeBodyVLA (the standout true loco-manip VLA), Lang-To-Loco, HWC-Loco, HVD, LIFT.
  • CVPR 2026 β€” the sim-to-real data engine. VIRAL, Open-Sim-to-Real, Gallant, Humanoid-GPT, EgoVLA β€” the vision venue owns how you generate humanoid training signal at scale (teacher-student, voxel terrain, mocap pretraining, ego-video).
  • Industry (no papers) β€” dual-system + teleop. Figure Helix and NVIDIA GR00T define the deployed template; Helix is upper-body, GR00T is cross-embodiment with a real-teleop core.
  • IROS 2026 πŸ†• β€” bimanual coordination structure + whole-body unification. The frontier is coordination architecture, not a second arm: 3D FlowMatch Actor (CMU/NVIDIA β€” one 3D policy for single and dual-arm, +41.4% PerAct2, beats 1000Γ—-larger models) and EquiBim (bilateral symmetry-equivariance as an inductive bias). Whole-body: ULTRA (unified multimodal whole-body loco-manip), CEER (compliant EE+root unified interface), OmniDP (beyond-FOV omnidirectional 3D perception), DreamMimic (loco-manip via a world model), and a MoE-VLA for humanoid loco-manip (Responsibility-Induced Specialized Experts β€” the multi-task interference fix reaching whole-body). Data is the bottleneck β†’ RL-based bimanual data generation + physics-informed retargeting (SPIDER). Context: IROS 2026 survey Β§5.3.

8. Comparison table

System Venue arXiv/ID Platform Β· DoF Axis 1 β€” Mobility/Balance Axis 2 β€” Action space Axis 3 β€” Data
LeVERB arXiv Jun 2025 2506.13751 G1 Β· ~23–29 RL WBC decodes latent β†’ walk/turn/sit Latent action vocabulary (A) Synthetic only; LeVERB-Bench 58.5%
WholeBodyVLA ICLR 2026 2512.11047 AgiBot X2 Β· bipedal LMO RL policy (advance/turn/squat) Unified latent + dual-stream head (A) Action-free ego video + thin teleop; +21.3%
HumanVLA NeurIPS 2024 2406.19972 Sim humanoid AMP teacher → balanced walk-carry Single whole-body policy Sim teacher→VLA student; Human-in-the-Room set
Humanoid-VLA arXiv Feb 2025 2502.14795 Humanoid (unspec) Whole-body control backbone β€” Language-motion align + ego video + pseudo-labels
GR00T N1–N1.7 NVIDIA 2503.14734 (N1) GR-1/G1/Atlas… N1 manipulation-only; G1 WBC in later code Dual-system DiT + per-embodiment MLP (B) Data pyramid (teleop+sim+ego); N1.7 = 20,854 h EgoScale
Figure Helix blog only no paper Figure 02 Β· 35 (upper) None (upper-body only) Dual-system 7B@7–9 Hz + 80M@200 Hz (B) ~500 h teleop (vendor)
DuoCore-FS arXiv Dec 2025 2512.20188 Astribot S1 Β· 25 Omitted (quasi-static) RVQ stream tokenizer (C); async 2.6Γ— 10.22 h teleop, 1 task
Galaxea G0 arXiv Sep 2025 2509.00576 Mobile base Mobile (wheeled, not bipedal) Dual-system planner+executor (B) Open dataset; 3-stage curriculum
BFM-Zero ICLR 2026 2511.04131 G1 · 29 Promptable FB controller; push recovery FB latent z∈R²⁡⁢ (E) LAFAN1 mocap; 192 M steps
Lang-To-Loco ICLR 2026 OR k3Cyx3Uets G1 Β· 23 Diffusion student over MoE teacher 64-D motion latent (E); no vision MotionMillion (retargeting-free)
HWC-Loco ICLR 2026 OR 3UE3Aatcjy H1/G1 Β· 19 ZMP-constrained robust locomotion 19-DoF, VAE privileged-state CMU mocap; sim-only
UniFP CoRL 2025 β˜… 2505.20829 B2-Z1/G1 Β· 18/29 Unified position+force loco-manip Force as first-class output Sim cmd combos; +39.5%
VideoMimic CoRL 2025 β˜… 2505.03729 G1 Β· 23 Stairs/terrain/sit-stand via root cmds Drop target-angle conditioning Human video β†’ 4D reconstruct β†’ retarget
VIRAL CVPR 2026 2511.15200 G1 Zero-shot loco-manip Delta actions + RSI Teacher→RGB student, 64 GPUs; 54 cycles
Open-Sim-to-Real CVPR 2026 2512.01061 Humanoid Door-opening loco-manip Pure RGB→action Staged-reset teacher → GRPO; 31.7% > human
EgoVLA CVPR 2026 2507.12440 H1 + 2Γ—12 hands Bimanual loco-manip Wrist SE(3) + MANO 500 K ego pairs; IK retarget; +43 pp long-horizon

β˜… = Best/Best-Student Paper. "OR" = OpenReview ID (no arXiv listed).


9. Decision guide

  1. Need true bipedal loco-manipulation (walk + manipulate)? β†’ F1 latent-vocabulary VLA over an RL WBC (WholeBodyVLA, LeVERB). Accept that it's mostly sim-trained today.
  2. Stationary/upper-body humanoid, want best manipulation? β†’ F2 dual-system (GR00T, Helix-style). Balance is not your problem; invest in the fast head and teleop.
  3. Balance/locomotion is the hard part, semantics secondary? β†’ F3 controller substrate first (BFM-Zero, HWC-Loco, UniFP), then prompt it with a VLM.
  4. Contact-rich loco-manip (push doors, wipe, carry)? β†’ unified position+force (UniFP) under the WBC.
  5. Data-constrained (no teleop farm)? β†’ human-video / sim engine: VideoMimic or EgoScale-style ego-video for manipulation; VIRAL / Open-Sim-to-Real teacher-student for vision sim-to-real.
  6. Can't hit real-time with a big VLM? β†’ async fast-slow (DuoCore-FS, Fast-in-Slow) and a high-rate System-1 head.

10. Limitations & open problems

  1. VLA-driven balance is still delegated, not learned end-to-end. Every credible loco-manip VLA puts an RL WBC underneath; nobody robustly emits balance-critical whole-body torques directly from a language-conditioned policy. Whether end-to-end is even desirable is open.
  2. The bipedal loco-manip VLA set is tiny and mostly simulated. LeVERB, WholeBodyVLA, HumanVLA β€” small sample, sim-heavy, few real-world long-horizon demos. Most "humanoid VLAs" are upper-body or mobile-base.
  3. Data is the binding constraint, and retargeting is fragile. Human-video transfer hinges on metric-scale reconstruction + IK under joint limits; EgoVLA shows human video alone gives 0% zero-shot. No standard humanoid-VLA benchmark exists (LeVERB-Bench, Isaac Humanoid Manip Benchmark, RoboCasa are not unified).
  4. Marketing β‰  capability. Figure Helix has no paper and is upper-body; vendor numbers are unverifiable. GR00T's later versions (N1.5+) are blog/model-card only, not peer-reviewed.
  5. High-DoF credit assignment is unsolved. HVD's value-decomposition helps offline RL but the broad problem β€” learning coordinated 30-DoF whole-body actions from limited data β€” remains hard.
  6. Compute walls. VIRAL shows vision sim-to-real for humanoids needs ~64 GPUs; this gates academic reproduction.
  7. Frontier, unverified (cite with care): 2026 preprints surfaced but not source-verified here β€” PhysiFlow (2603.05410, multi-brain latent flow-matching whole-body VLA), HEX (2604.07993, humanoid-aligned experts for cross-embodiment whole-body), HumanoidExo (2510.03022, exoskeleton-data whole-body VLA), Cybo-Waiter (2603.10675). Verify before relying on them.

11. Adjacent systems (not bipedal β€” included to prevent miscategorization)


12. Links


πŸ—“ State of the Field (updated Aug 2026)

Verdict: the triple-system recipe (VLM + flow expert + RL lower body) became the open reference at RSS 2026; data efficiency, not data volume, is the winning argument β€” and evaluation lags arms by a generation.

πŸ“ˆ Trend

RSS 2026 was humanoid loco-manipulation's coming-out: Ξ¨β‚€ (open foundation model; 800 h video + 30 h robot data beats >10Γ— co-trained corpora incl. GR00T N1.6 by >40 pp), HiWET (world-frame EE tracking beats body-frame), HAIC (dynamics-aware interaction WM), EgoHumanoid + HoMMI (robot-free demonstration pipelines), plus a deep whole-body-control bench (TeleGate, OmniXtreme, X-Loco, pixel locomotion).

βš–οΈ Approaches & trade-offs

Fork Options Trade-off
Data recipe Co-train human+humanoid in one policy (GR00T/EgoVLA/H-RDT line) vs decouple video→representation / robot→control (Ψ₀) Hardware evidence favors decoupling at 36-DoF scale — the humanoid face of the human-video fork
Collection interface Decoupled teleop rigs (PICO+MANUS+trackers; locomotion delegated) vs robot-free human interfaces (HoMMI's UMI+ego) Stability & fidelity vs scalability
Control stack Whole-body end-to-end vs triple-system (System-0 RL lower body) Expressiveness vs stability; RL controller caps agility (no dynamic bracing yet)

⚠️ Limitations & open problems

  • Per-task fine-tuning sits inside every published loop (Ξ¨β‚€: 80 demos/task) β€” no humanoid zero-shot generalist.
  • Payload and hand-DoF (Dex3-1-class) cap difficulty; precision insertion unsolved.
  • No humanoid OOD protocol or perturbation suite; intervention-assisted scoring is the norm.

Latest (preprint): Ο‰-0 πŸ†• β€” whole-body latent-predictive WAM for concurrent loco-manipulation (81.8% on 11 household tasks vs 44.5% ψ-0). See Latest Papers.

← Back to Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally