-
Notifications
You must be signed in to change notification settings - Fork 0
Review RLDX 1
Model: RLDX-1 β dexterity-first robot-hand foundation model Β· RLWRLD (Seoul, South Korea) Announced: May 2026 (SF Exploratorium launch) Β· built on NVIDIA's physical-AI stack. Status: industry release β all figures below are company/vendor claims, not peer-reviewed. Filed alongside DYNA-2 as an industrial datapoint. Sources: RLWRLD Β· The Robot Report
β οΈ Sourcing caveat. RLDX-1 is a company launch; there is no peer-reviewed paper, no released weights, and the benchmarks are RLWRLD's own vs its chosen baselines. Treat every number as a vendor claim.
Companions: Dexterous-Hand Data Pyramid Β· VLA Hybrid Architectures Β· DYNA-2 Β· Dexterous Manipulation.
- Four modality streams in one transformer. RLDX-1 uses a Multi-Stream Action Transformer (MSAT) β Vision (fine-tuned Qwen3-VL 8B, spatial reasoning), Motion (spatio-temporal video features + object velocity), Memory (64 learnable "cognition tokens" for long-horizon), and Torque/Tactile (a Physics Module treating tactile + torque as native modalities for weight estimation and contact detection). "Each modality gets its own processing stream, and joint self-attention lets them interact."
- Human-hand-first data. Rather than teleop-heavy collection, RLDX records from the bare human hand and closes the gap in software with a retargeting framework built for five-finger dexterity β reportedly >200 demonstrations/hour.
- Synthetic augmentation. Video-generation models vary lighting/surface/position; inverse dynamics annotate actions; quality-filtered. A claimed ~5Γ data-scale increase β +9.2% average success.
- Runs across embodiments. WIRobotics ALLEX humanoid (five-finger hands), Franka Research 3 (AnySkin tactile + joint torque), OpenArm + Inspire 6-DoF hand.
- An industrial bet that mirrors the research pyramid. RLDX-1 explicitly composes bare-human-hand capture β five-finger retargeting β synthetic augmentation β small teleop set β the exact L1/L2 β L4 β L5 β L6 stack the data pyramid describes, with torque/tactile as a native stream (the cross-cut). It's the commercial instantiation of "human-hand data first, retargeting is the bridge."
- MSAT is a multi-stream MoT. Per-modality streams + joint self-attention is the same family as the three-expert Mixture-of-Transformers in Review-VLA-Hybrid-Architectures β here specialized to add motion, memory, and physics streams for contact-rich dexterity.
- But it's vendor-reported. Unlike the arXiv works on the pyramid, RLDX-1 has no paper/weights/independent eval; the Ο0.5 / GR00T comparisons are RLWRLD's own.
MSAT streams:
| Stream | What it does |
|---|---|
| Vision | robot-specialized VLM (fine-tuned Qwen3-VL 8B), spatial reasoning |
| Motion | spatio-temporal features from video; tracks object velocity |
| Memory | 64 learnable cognition tokens compress the scene for long-horizon tasks |
| Torque / Tactile | Physics Module β tactile + torque as native modalities β weight estimation, contact detection |
Data pipeline: bare-human-hand recording + five-finger kinematic retargeting (>200 demos/hr) β synthetic augmentation (video-gen trajectories + inverse-dynamics action labels + quality filtering) β small real teleop set. Built on NVIDIA's physical-AI stack.
Embodiments: ALLEX humanoid (five-finger) Β· Franka Research 3 (AnySkin + torque) Β· OpenArm + Inspire 6-DoF; supports single-arm / dual-arm / humanoid.
- ALLEX humanoid: baselines report "success rates below 30%"; RLDX-1 "reaches nearly 90%" on motion / history / physical-signal tasks.
- OpenArm: RLDX-1 keeps "balanced performance across task types," while GR00T N1.6 "completely fails on the object identification task."
- Conveyor pick-and-place: +37.5 percentage points over GR00T N1.6.
- Data scaling: ~5Γ data β +9.2% average success.
(No absolute protocol, held-out set, or independent replication is disclosed.)
Significance. RLDX-1 is a notable industrial vote for the human-hand-first + retargeting + synthetic data pyramid, and for torque/tactile as a first-class stream in a multi-stream transformer β aimed squarely at five-finger, contact-rich humanoid dexterity.
Limitations.
- Vendor-reported, not peer-reviewed / not replicated; no weights, no shared benchmark, self-chosen baselines.
- Bare-human-hand retargeting fidelity is the load-bearing assumption (the pyramid's L4 bottleneck) β asserted, not independently measured.
- Synthetic-augmentation quality (video-gen + inverse dynamics) inherits pixel-world-model limits on contact physics.
- Absolute success protocols undisclosed β "nearly 90%" / "below 30%" lack task-level detail.
- Sources: RLWRLD RLDX-1 Β· The Robot Report
- Pyramid placement: L1/L2 human-hand β L4 five-finger retarget β L5 synthetic β L6 teleop, + torque/tactile stream β Dexterous-Hand Data Pyramid
- Architecture kin: VLA Hybrid Architectures (multi-stream MoT) Β· industrial sibling: DYNA-2
- Dexterous Manipulation Β· Tactile VLA
- Home
- π Changelog
- πΈοΈ Knowledge Graph
- π Latest Papers
- All in-depth reviews β topic catalog Β· per-paper
- VLA Architectures
- RL for VLA
- World Models
- Dexterous Manipulation
- Cross-Embodiment
- Humanoid VLA
(each page indexes its per-paper pages)