Skip to content

Review RLDX 1

hwoo.han edited this page Aug 13, 2026 · 1 revision

In-Depth Review β€” RLDX-1: A Dexterity-First Foundation Model for Robot Hands (RLWRLD)

Model: RLDX-1 β€” dexterity-first robot-hand foundation model Β· RLWRLD (Seoul, South Korea) Announced: May 2026 (SF Exploratorium launch) Β· built on NVIDIA's physical-AI stack. Status: industry release β€” all figures below are company/vendor claims, not peer-reviewed. Filed alongside DYNA-2 as an industrial datapoint. Sources: RLWRLD Β· The Robot Report

⚠️ Sourcing caveat. RLDX-1 is a company launch; there is no peer-reviewed paper, no released weights, and the benchmarks are RLWRLD's own vs its chosen baselines. Treat every number as a vendor claim.

Companions: Dexterous-Hand Data Pyramid Β· VLA Hybrid Architectures Β· DYNA-2 Β· Dexterous Manipulation.


1. TL;DR

  1. Four modality streams in one transformer. RLDX-1 uses a Multi-Stream Action Transformer (MSAT) β€” Vision (fine-tuned Qwen3-VL 8B, spatial reasoning), Motion (spatio-temporal video features + object velocity), Memory (64 learnable "cognition tokens" for long-horizon), and Torque/Tactile (a Physics Module treating tactile + torque as native modalities for weight estimation and contact detection). "Each modality gets its own processing stream, and joint self-attention lets them interact."
  2. Human-hand-first data. Rather than teleop-heavy collection, RLDX records from the bare human hand and closes the gap in software with a retargeting framework built for five-finger dexterity β€” reportedly >200 demonstrations/hour.
  3. Synthetic augmentation. Video-generation models vary lighting/surface/position; inverse dynamics annotate actions; quality-filtered. A claimed ~5Γ— data-scale increase β†’ +9.2% average success.
  4. Runs across embodiments. WIRobotics ALLEX humanoid (five-finger hands), Franka Research 3 (AnySkin tactile + joint torque), OpenArm + Inspire 6-DoF hand.

2. Why it matters (and the caveats)

  • An industrial bet that mirrors the research pyramid. RLDX-1 explicitly composes bare-human-hand capture β†’ five-finger retargeting β†’ synthetic augmentation β†’ small teleop set β€” the exact L1/L2 β†’ L4 β†’ L5 β†’ L6 stack the data pyramid describes, with torque/tactile as a native stream (the cross-cut). It's the commercial instantiation of "human-hand data first, retargeting is the bridge."
  • MSAT is a multi-stream MoT. Per-modality streams + joint self-attention is the same family as the three-expert Mixture-of-Transformers in Review-VLA-Hybrid-Architectures β€” here specialized to add motion, memory, and physics streams for contact-rich dexterity.
  • But it's vendor-reported. Unlike the arXiv works on the pyramid, RLDX-1 has no paper/weights/independent eval; the Ο€0.5 / GR00T comparisons are RLWRLD's own.

3. Architecture & data (as described)

MSAT streams:

Stream What it does
Vision robot-specialized VLM (fine-tuned Qwen3-VL 8B), spatial reasoning
Motion spatio-temporal features from video; tracks object velocity
Memory 64 learnable cognition tokens compress the scene for long-horizon tasks
Torque / Tactile Physics Module β€” tactile + torque as native modalities β†’ weight estimation, contact detection

Data pipeline: bare-human-hand recording + five-finger kinematic retargeting (>200 demos/hr) β†’ synthetic augmentation (video-gen trajectories + inverse-dynamics action labels + quality filtering) β†’ small real teleop set. Built on NVIDIA's physical-AI stack.

Embodiments: ALLEX humanoid (five-finger) Β· Franka Research 3 (AnySkin + torque) Β· OpenArm + Inspire 6-DoF; supports single-arm / dual-arm / humanoid.


4. Reported results (⚠️ vendor claims vs Ο€0.5 & GR00T N1.6)

  • ALLEX humanoid: baselines report "success rates below 30%"; RLDX-1 "reaches nearly 90%" on motion / history / physical-signal tasks.
  • OpenArm: RLDX-1 keeps "balanced performance across task types," while GR00T N1.6 "completely fails on the object identification task."
  • Conveyor pick-and-place: +37.5 percentage points over GR00T N1.6.
  • Data scaling: ~5Γ— data β†’ +9.2% average success.

(No absolute protocol, held-out set, or independent replication is disclosed.)


5. Significance & limitations

Significance. RLDX-1 is a notable industrial vote for the human-hand-first + retargeting + synthetic data pyramid, and for torque/tactile as a first-class stream in a multi-stream transformer β€” aimed squarely at five-finger, contact-rich humanoid dexterity.

Limitations.

  1. Vendor-reported, not peer-reviewed / not replicated; no weights, no shared benchmark, self-chosen baselines.
  2. Bare-human-hand retargeting fidelity is the load-bearing assumption (the pyramid's L4 bottleneck) β€” asserted, not independently measured.
  3. Synthetic-augmentation quality (video-gen + inverse dynamics) inherits pixel-world-model limits on contact physics.
  4. Absolute success protocols undisclosed β€” "nearly 90%" / "below 30%" lack task-level detail.

6. Links

← Back to Reviews Β· Home

Navigation

πŸ“– Reviews

🏷 Model lineages

🧠 ML foundations

πŸ—“ Conferences

(each page indexes its per-paper pages)

πŸ“Œ Foundational

Clone this wiki locally