Skip to content

Releases: Janga786/xembench

v0.5.0 — Phase B results, data flywheel round 1, precision-manipulation study

Choose a tag to compare

@Janga786 Janga786 released this 11 Jul 05:26

First results-bearing release. Everything traces to checkpoint + config + git SHA.

Headline numbers (73-cell matrix, 3 seeds, Wilson CIs, corrected horizons)

  • push_t (Panda stick): 45.7% BC native success from pixels+language; robust to held-out paraphrases (46.3%)
  • transport_box (Unitree G1 humanoid): 18.8% — learned from stochastic-expert demonstrations
  • pull_tool 1.3% → 14.7% with action chunking + 10× demos (matrix-grade: 3 seeds × 600 episodes, held-out paraphrase 15.3%)
  • Precision tasks (pick/stack/apple): 0% with the cause isolated — see the intervention study in the README (pipeline exonerated by 5/5 ground-truth replay; chunking, ensembling, 50× data, encoder capacity, and RL fine-tuning systematically eliminated; the policy class is the quantified open problem)

Also in this release

  • Physical AI Data Flywheel round 1 executed on real data (targeted vs random arms, CI-gated comparison, first non-synthetic flywheel pack)
  • Action chunking (ACT-style) with queue/ensemble/receding-horizon execution, backward compatible
  • 12,900 validated demonstrations (1,202 benchmark set + scaling sets)
  • Rollout recording hook re-running exact evaluation episode seeds; media in README
  • Six root-caused engineering findings (PhysX collision-stack growth, metric window bias, KL detonation under staged rewards, mirrored G1 hand URDF, …)
  • MIT license, CI (fixed to actually run — it never had), campaign driver scripts under scripts/campaign/

Full account: reports/phase_b_campaign_report.md