-
Notifications
You must be signed in to change notification settings - Fork 0
RSS 2026 TeleGate
Venue: RSS 2026 (Sydney, Jul 13β17) Β· Session: Humanoids Β· paper #25 Authors: Jie Li, Bing Tang, Feng Wu arXiv: 2602.09628 Β· program page
Summary compiled from the arXiv paper (v2, USTC + AnyWit Robotics); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 shows an operator in an inertial motion-capture suit teleoperating the Unitree G1 in real time across four highly dynamic behaviors: (a) grasping and placing toys into a basket, (b) standing long jump, (c) prone-position stand-up, and (d) kicking a ball β the motion classes that typically break distilled single-policy controllers.
Real-time whole-body humanoid teleoperation needs one controller that handles motions with very different dynamics (walking, jumping, fall recovery), but multi-motion RL suffers catastrophic forgetting and the standard fix β distilling multiple expert policies into one generalist β loses performance on highly dynamic motions. Real-time teleoperation additionally lacks the future reference trajectory that offline trackers exploit, so anticipatory motions (jumping, standing up) are hard.
TeleGate has three stages: (I) collect whole-body motion with inertial mocap and retarget it online to the robot, partitioning data into six motion categories; (II) train four expert policies (walk/run, dance/martial-arts, fall-and-recovery, jump) with PPO in MuJoCo using asymmetric actorβcritic, plus a jointly trained small-Transformer VAE motion-prediction prior whose encoder reads the historical trajectory window and whose decoder reconstructs the future window, so the latent z_t injects implicit future-motion intent into the policy (loss = PPO + reconstruction + KL); a failure-rate-based curriculum reweights hard clips; (III) freeze all experts and train a lightweight gating network that scores the K experts from proprioception and reference trajectory and routes each control cycle to the top-1 expert. Training uses only ~2.5 hours of mocap (walking 40 min, running 24, dancing 24, martial arts 20, fall recovery 26, jumping 16).
In 4-seed simulation comparisons TeleGate reaches 97.3% success and 17.22 mm mean per-joint position error, versus TWIST 68.9%/53.55 mm, Any2track 91.2%/18.40 mm, and GMT 92.0%/29.14 mm. Ablations show top-1 gating (96.7%) beats DAgger distillation (91.2%), a single policy (94.7%), top-2 gating (93.7%), and random expert partitioning (95.0%), and even edges the per-expert Oracle (96.3%); adding the VAE prior lifts overall SR to 97.3% and cuts Jump Empjpe by 13% (22.27β19.40 mm) and Fall-Recovery by 8% (28.23β25.94 mm). The system is deployed on a real Unitree G1 performing running, standing long jump, prone stand-up, and ball kicking.
Shows that routing among frozen experts, rather than distilling them, preserves expert-level dynamic skill in a unified real-time teleoperation controller with tiny data (2.5 h vs 42β700 h for TWIST/SONIC-scale efforts). Relevant to the humanoid control stack discussion in Review-Humanoid-VLA and the fast/slow control decomposition in Review-System-0-1-2.
β Back to RSS 2026 survey Β· RSS-2026-Papers Β· Home
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)