Skip to content

CLEAR-LeWM v0.3.0: task-semantic evaluation

Latest

Choose a tag to compare

@DavidSunok DavidSunok released this 22 Jul 10:44

CLEAR-LeWM v0.3.0 makes benchmark success correspond to physical task completion across all four LeWM tasks.

Highlights:

  • Two auditable protocols: Moderate for continuity and Strict for stronger task completion claims
  • PushT object-only pose evaluation that excludes the pusher endpoint
  • Cube position plus symmetry-aware orientation over all 24 proper cube rotations
  • Reacher wrapped joint geometry with holding and terminal speed reported separately
  • TwoRoom swept-circle collision, full-radius door clearance, and mandatory legal-route validation
  • Deterministic versioned manifests with matched random-policy controls
  • Historical Official, Moderate, and Strict outputs with checkpoint and runtime provenance
  • Official high-epoch LeWM checkpoint registry and strict checkpoint preflight
  • Rebuilt 1080p rule-comparison film, high-resolution task GIFs, and task-specific evaluation guides

Reference results (LeWM / random, 100 pairs, seed 42, 300 x 30 CEM, solver batch 1):

  • PushT Strict: 79% / 2%
  • Cube Strict: 18% / 3%
  • Reacher Strict: 36% / 4%
  • TwoRoom Strict: 24% / 0%, with 0/100 invalid routes

Compatibility:

  • v0.1 and v0.2 artifacts remain executable and are not silently rewritten.
  • v0.3 protocols, manifests, checkpoint registry, and results are versioned explicitly.

Validation:

  • 37 tests passed; 2 optional-dependency tests skipped
  • Ruff lint and repository-wide formatting checks passed
  • README local links validated
  • All release MP4/GIF media fully decoded successfully