First results-bearing release. Everything traces to checkpoint + config + git SHA.
Headline numbers (73-cell matrix, 3 seeds, Wilson CIs, corrected horizons)
- push_t (Panda stick): 45.7% BC native success from pixels+language; robust to held-out paraphrases (46.3%)
- transport_box (Unitree G1 humanoid): 18.8% — learned from stochastic-expert demonstrations
- pull_tool 1.3% → 14.7% with action chunking + 10× demos (matrix-grade: 3 seeds × 600 episodes, held-out paraphrase 15.3%)
- Precision tasks (pick/stack/apple): 0% with the cause isolated — see the intervention study in the README (pipeline exonerated by 5/5 ground-truth replay; chunking, ensembling, 50× data, encoder capacity, and RL fine-tuning systematically eliminated; the policy class is the quantified open problem)
Also in this release
- Physical AI Data Flywheel round 1 executed on real data (targeted vs random arms, CI-gated comparison, first non-synthetic flywheel pack)
- Action chunking (ACT-style) with queue/ensemble/receding-horizon execution, backward compatible
- 12,900 validated demonstrations (1,202 benchmark set + scaling sets)
- Rollout recording hook re-running exact evaluation episode seeds; media in README
- Six root-caused engineering findings (PhysX collision-stack growth, metric window bias, KL detonation under staged rewards, mirrored G1 hand URDF, …)
- MIT license, CI (fixed to actually run — it never had), campaign driver scripts under scripts/campaign/
Full account: reports/phase_b_campaign_report.md