Skip to content

CVPR 2026 Humanoid GPT

Heungwoo edited this page Jun 1, 2026 · 2 revisions

Humanoid-GPT β€” Humanoid Generative Pre-Training for Zero-Shot Motion Tracking

Venue: CVPR 2026 (Poster #39399) Category: Humanoid Motion / Pretraining Trend tag: Trend 2 Affiliations: Tsinghua University Β· Galbot Β· Beihang University Β· Shanghai Jiao Tong University Β· Peking University Β· Shanghai Qi Zhi Institute

Approach diagram

flowchart LR
  MOCAP["2B-frame retargeted mocap corpus<br/>AMASS Β· LAFAN1 Β· Motion-X++ Β· PHUMA Β· MotionMillion + in-house"] --> HME["HME diversity clustering"]
  HME --> EXP["per-cluster RL motion experts"]
  EXP --> DAGGER["DAgger distillation"]
  DAGGER --> GPT["single causal-attention GPT<br/>generalist tracker"]
  TARGET["target trajectory"] --> GPT
  GPT --> ACT["zero-shot humanoid action sequence"]
Loading

Problem

Humanoid motion-tracking policies have always been task-specialist: train per task, retrain for each new motion. The agility-vs-generalization trade-off is fundamental in conventional designs.

Method

Humanoid-GPT is not a single GPT trained directly on raw mocap. The pipeline is expert-distillation:

  1. Corpus + diversity metric. A 2 B-frame retargeted corpus unifies all major mocap datasets (AMASS, LAFAN1, Motion-X++, PHUMA, MotionMillion) with large-scale in-house recordings β€” ~4–5Γ— larger log-volume than AMASS under the authors' Harmonic Motion Embedding (HME), a novel metric that quantifies and categorizes motion diversity directly from motion data.
  2. Per-cluster RL experts. Motion experts are trained via reinforcement learning on HME-clustered data.
  3. DAgger distillation. A GPT-style Transformer with causal temporal attention is trained via DAgger to consolidate all expert controllers into one generalist tracker.

At deployment, conditioning on a target trajectory yields a zero-shot motion tracker that runs on a Unitree G1.

Results

  • Zero-shot generalization to unseen motions and control tasks (e.g., unseen dance sequences) while still tracking highly dynamic in-domain motions β€” directly attacking the agility-vs-generalization trade-off.
  • Clear scaling laws: enlarging both corpus and model capacity yields consistent gains in tracking accuracy and stability.
  • Latency: end-to-end inference under 1.5 ms on a single NVIDIA RTX 4090 (ONNX + TensorRT), ~5Γ— faster than TWIST.

Significance

Humanoid-GPT shows that a single causal Transformer, distilled from many per-cluster RL experts over a billion-scale, diversity-balanced corpus, can beat per-task specialist trackers and generalize zero-shot. The contribution is two-fold: (i) the HME-driven data pipeline that makes "scale" meaningful by measuring and balancing motion diversity, and (ii) the DAgger distillation that turns a fleet of RL experts into one generalist. Its lineage is the whole-body motion-tracking line (it benchmarks against TWIST), not VLA manipulation.

Links

Related pages

← Back to CVPR-2026

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally