Predictability cliffs are intrinsic structures of long-horizon manipulation tasks. This repository implements the cliff detection, closed-loop replanning, and phenomenon discovery experiments from the PACE v2 paper (CoRL 2027 submission).
Read this table before any number in this README. Anything tagged
placeholderis a synthetic dry-run output; onlyverifieditems are checked numerically against the math.
| Item | Status | Notes |
|---|---|---|
| Phase InfoNCE identifiability ( |
verified (CPU) |
verify_identifiability.py PA |
|
|
verified (CPU) |
verify_phase_posterior.py peak-F1 |
| PCAR budget-quantile DKW bound | verified (CPU) |
verify_pcar_budget.py |
| Boundary-aware flow loss reduction | verified (CPU) |
sanity_pace_a.py |
| Unit + smoke tests | verified (CPU) |
pytest tests/ → 212 passed (v2.1) |
|
|
verified (CPU) |
compute_I_hat_2 implemented; tests in test_cliff_estimators.py
|
|
|
verified (CPU) |
compute_I_hat_3 implemented; cached c_{t-1} cleared by policy.reset()
|
| Concordance |
verified (CPU) |
compute_concordance_C rolling-window rank fusion; tests cover [0,1] range and online state |
| LIBERO-Long / Spatial SR (Tables 1, 2) | placeholder | requires GPU + LIBERO dataset + trained checkpoint |
| Phenomenon §6.1, §6.4 numbers | placeholder | dry-run pipeline; full results require real env + checkpoints |
| Inference cost (params, NFE, latency) | placeholder | parameter count and NFE are correct per architecture; latency requires real GPU benchmark |
This environment is CPU-only with no LIBERO/SimplerEnv installed and
no trained checkpoints. The CPU-verifiable items above are what
this repo can prove correct today; everything tagged placeholder
needs a GPU + datasets to produce real numbers (see
Verification & Reproduction Guide below).
# 1. Install (CPU is enough for verification scripts and unit tests)
conda env create -f environment.yml && conda activate lerobot_env
pip install -r requirements.txt
pip install -e ./lerobot_policy_phaseqflow
# 2. Smoke test (no GPU, no data required)
bash scripts/smoke/smoke_phase_centric.sh
# 3. Run all CPU-only verifications (~3 min total)
pytest tests/ -q
python scripts/verification/verify_identifiability.py
python scripts/verification/verify_phase_posterior.py
python scripts/verification/sanity_pace_a.py
python scripts/verification/verify_pcar_budget.pySee docs/OPERATIONS_GUIDE.md for the engineering handbook and docs/ARCHITECTURE.md for the full specification, including the math behind each verification script.
Status: GPU evaluation not yet run. Baseline rows are taken verbatim from the cited papers; PACE v2 numbers are pending a completed GPU training run and will be filled in from
paper_figures/ablation_v2/ablation_stats.csv.
| Method | LIBERO-Long SR | Source |
|---|---|---|
| OpenVLA-OFT | 54.5% | arXiv 2502.19645, Table 2 |
| π₀ | 60.0% | arXiv 2410.24164, Table 1 |
| MoH | 57.8% | arXiv 2410.11842, Table 3 |
| PACE v2 config 07 (ours) | pending GPU run | — |
| Oracle cliff (config 06, upper bound) | pending GPU run | — |
Note: cited papers may use different LIBERO-Long splits or evaluation protocols. Direct numeric comparison requires a shared evaluation harness.
Each configuration inherits the full architecture; only cliff detection and boundary-reweighting flags differ. CoRL submission uses three configs × 3 seeds, evaluated with IQM ± 95% bootstrap CI (rliable-style):
Status: all numbers below are synthetic dry-run placeholders generated by
scripts/aggregate_ablation.py --dry_run. Real numbers land after the cloud sweep finishes (scripts/run_autodl_pipeline.sh train).
v2.1 implementation status:
compute_I_hat_1(Bhattacharyya),compute_I_hat_2(action variance),compute_I_hat_3(velocity curvature), andcompute_concordance_C(rank-window fusion) are all implemented and exercised bytests/test_cliff_estimators.py. The CoRL submission's headline ablation 07 uses all three estimators internally via the concordance signal, so the trimmed table still validates the full theoretical chain.
| Config | Description | LIBERO-Long IQM (placeholder) | LIBERO-Spatial IQM (placeholder) |
|---|---|---|---|
| 01 | BC-Chunked (baseline) | 0.520 | 0.634 |
| 02 | Cliff via β̂_t only (I^(1)) | 0.593 | 0.690 |
| 07 | PACE v2: C_t + boundary reweight | 0.692 | 0.727 |
All numbers above are synthetic dry-run placeholders pending the full LIBERO-Long / LIBERO-Spatial GPU sweep.
# Verify pipeline (synthetic data, no checkpoint):
python scripts/aggregate_ablation.py --dry_run
# Real aggregation after GPU training:
python scripts/aggregate_ablation.py \
--input_root outputs/ablation_v2 \
--output paper_figures/ablation_v2/All numbers in §6.1–6.5 are placeholder outputs from
--dry_runmode (synthetic distributions hard-coded inside each script). They verify the analysis pipeline but do not reflect real LIBERO rollouts. Real numbers are produced by running each script without--dry_runagainst trained checkpoints; see the Verification & Reproduction Guide.
The pipeline rolls out OpenVLA-7B, π₀, BC-ACT, and Diffusion Policy on
LIBERO-Long, runs cliff detection, and measures the time gap between the
last detected cliff and the first failure step. The hypothesis is that
the gap distribution is policy-agnostic (right-skewed, mass within
5–25 steps of the last cliff). Pairwise KS test
# Placeholder (CPU, synthetic):
python scripts/phenomenon/universality.py --dry_run
# Real (requires the four checkpoints + LIBERO-Long environment):
python scripts/phenomenon/universality.py \
--n_rollouts 50 --seeds 0 1 2 \
--output paper_figures/universality/Hypothesis: success-rate regret
| Chunk length |
SR (placeholder) | $\delta$SR (placeholder) |
|---|---|---|
| 4 | 0.829 | 0.000 |
| 8 | 0.846 | 0.000 |
| 16 | 0.712 | 0.007 |
| 32 | 0.814 | 0.025 |
| 64 | 0.718 | 0.048 |
Linear fit through origin reports
# Placeholder (CPU, synthetic):
python scripts/phenomenon/regret_scaling.py --dry_run
# Real (requires PACE v2 checkpoint + LIBERO-Long):
python scripts/phenomenon/regret_scaling.py \
--checkpoint checkpoints/pace_v2_libero_long \
--H_values 4 8 16 32 64 \
--output paper_figures/regret_scaling/Precision/recall of each estimator and the concordance
| Detector | Precision | Recall | F1 |
|---|---|---|---|
|
|
0.162 | 1.000 | 0.279 |
|
|
0.155 | 1.000 | 0.269 |
|
|
0.153 | 1.000 | 0.266 |
| Concordance |
1.000 | 1.000 | 1.000 |
All numbers above are placeholders. The trivial 1.000/1.000/1.000 row
for
python scripts/phenomenon/triangulation_concordance.py --dry_runComparison of six detector strategies on a gripper-flip oracle,
| Detector | Precision | Recall | F1 |
|---|---|---|---|
| Concordance |
0.429 | 1.000 | 0.600 |
| Bhattacharyya |
0.353 | 1.000 | 0.522 |
| KL divergence |
0.347 | 1.000 | 0.515 |
| JS divergence |
0.353 | 1.000 | 0.522 |
| Posterior entropy |
0.353 | 0.973 | 0.519 |
| BOCPD | 0.071 | 0.953 | 0.133 |
python scripts/diagnostics/diagnostic_utils/trigger_comparison.py --dry_runHypothesis: the flow-matching loss at phase-boundary steps is
materially higher than at interior steps, motivating boundary-aware
reweighting. The placeholder generator hard-codes
$\mathbb{E}{\text{boundary}} = 0.20$, $\mathbb{E}{\text{interior}} = 0.05$,
producing a ratio of
python scripts/diagnostics/diagnostic_utils/measure_boundary_error.py --dry_runParameter counts and NFE per inference path are derived from the architecture (correct by construction); the latency and Hz columns require a real GPU benchmark and are placeholders.
| Method | Params | NFE | Latency (placeholder) | Hz (placeholder) |
|---|---|---|---|---|
| BC-Chunked | 312.4M | 4 | 12.4 ms | 81 |
| Cliff via |
312.5M | 4 | 12.7 ms | 79 |
| Cliff via |
312.5M | 36 | 91.4 ms | 11 |
| Cliff via |
312.5M | 8 | 23.3 ms | 43 |
| Concordance |
312.6M | 44 | 96.5 ms | 10 |
| PACE v2 ( |
312.6M | 44 | 99.1 ms | 10 |
NFE derivation:
# Placeholder latency (CPU; the latency column is hardcoded synthetic):
python scripts/diagnostics/diagnostic_utils/measure_inference_cost.py --dry_run
# Real (requires GPU + checkpoint):
python scripts/diagnostics/diagnostic_utils/measure_inference_cost.py \
--checkpoint checkpoints/pace_v2_libero_long \
--device cuda --batch_size 1This section is the complete recipe to take this repo from a clean checkout to the figures and tables in the paper. It is split into five parts in order of increasing resource demand:
-
Part A: CPU-only verifications of the math (no GPU, no data,
$\le 5$ min) - Part B: Prerequisites for real experiments (GPU, datasets, checkpoints)
- Part C: Three-stage training pipeline
- Part D: Evaluation pipeline (LIBERO-Long, LIBERO-Spatial, SimplerEnv)
- Part E: Phenomenon experiments and figure generation
These verify mathematical claims on synthetic distributions and are the gate for any code change.
# A.1 Environment
conda env create -f environment.yml && conda activate lerobot_env
pip install -r requirements.txt
pip install -e ./lerobot_policy_phaseqflow
# A.2 Unit + smoke tests (~30 s)
pytest tests/ -q # expects: 204 passed
bash scripts/smoke/smoke_phase_centric.sh # 7-mode E2E, all [OK]
# A.3 Math verifications (~3 min)
python scripts/verification/verify_identifiability.py # PA >= 0.7 (or WARN_DEGENERATE on CPU)
python scripts/verification/verify_phase_posterior.py # peak-F1 >= 0.5
python scripts/verification/sanity_pace_a.py # boundary FM drop >= 20%
python scripts/verification/verify_pcar_budget.py # |rate - eps| < 0.005
# A.4 Pipeline dry-runs (verifies orchestration; numbers are synthetic)
python scripts/aggregate_ablation.py --dry_run
python scripts/phenomenon/universality.py --dry_run
python scripts/phenomenon/regret_scaling.py --dry_run
python scripts/phenomenon/triangulation_concordance.py --dry_run
python scripts/diagnostics/diagnostic_utils/trigger_comparison.py --dry_run
python scripts/diagnostics/diagnostic_utils/measure_boundary_error.py --dry_run
python scripts/diagnostics/diagnostic_utils/measure_inference_cost.py --dry_runIf every line in A.2–A.3 returns PASS / [OK] / 204 passed, the
implemented math is correct. Part A does not produce paper numbers.
| Resource | Requirement | Where to get it |
|---|---|---|
| GPU |
|
local or cloud |
| LIBERO datasets | LIBERO-10 (long-horizon) and LIBERO-Spatial | https://libero-project.github.io |
| SimplerEnv | for §6 universality on RT-1 / OpenVLA simulator | https://github.com/simpler-env/SimplerEnv |
| Baseline checkpoints | OpenVLA-7B, π₀, BC-ACT, Diffusion Policy | HuggingFace; loaded via baselines/*_adapter.py
|
| Storage |
|
— |
The four baseline adapters in baselines/ already handle download +
load via is_available() / load() / rollout(); populate
HF_HOME and the adapter load() will fetch the weights on first
call.
Once datasets are placed under data/libero/ and checkpoints under
checkpoints/, the pipeline below runs without further manual fetch.
Stage 1–3 produce a single PACE v2 checkpoint. Each stage is a
separate config under configs/train/.
# Stage 1: pretrain multimodal tokenizer (~2 days on RTX 5070)
python scripts/training/train_dummy_batch.py \
--config configs/train/01_pretrain_multimodal.yaml \
--total_steps 50000 --device cuda \
--output checkpoints/stage1_pretrain/
# Stage 2: hierarchical phase encoder + boundary-aware flow head (~3 days)
python scripts/training/train_dummy_batch.py \
--config configs/train/02_train_phase_and_flow.yaml \
--resume checkpoints/stage1_pretrain/last.pt \
--total_steps 100000 --device cuda \
--output checkpoints/stage2_phase_flow/
# Stage 3: freeze main net, calibrate PCAR / B-PCAR / concordance (~6 hours)
python scripts/training/train_dummy_batch.py \
--config configs/train/03_finetune_replan.yaml \
--resume checkpoints/stage2_phase_flow/last.pt \
--total_steps 5000 --device cuda \
--output checkpoints/pace_v2_libero_long/
# Calibration helpers (run after stage 3):
python scripts/calibration/calibrate_concordance.py \
--checkpoint checkpoints/pace_v2_libero_long \
--output checkpoints/pace_v2_libero_long/concordance_threshold.json
python scripts/calibration/calibrate_b_pcar.py \
--checkpoint checkpoints/pace_v2_libero_long \
--output checkpoints/pace_v2_libero_long/b_pcar_params.jsonEvaluates the trained checkpoint on LIBERO-Long, LIBERO-Spatial,
and SimplerEnv. Each rollout records SR + PCAR replan timings to
eval_results.json.
# D.1 LIBERO-Long (50 episodes × 3 seeds)
bash scripts/eval/run_libero_main.sh \
--checkpoint checkpoints/pace_v2_libero_long \
--dataset LIBERO-10 \
--seeds 42 123 2024 \
--output outputs/eval/libero_long/
# D.2 LIBERO-Spatial (specificity test, §4.4 of ARCHITECTURE.md)
bash scripts/eval/run_libero_main.sh \
--checkpoint checkpoints/pace_v2_libero_long \
--dataset LIBERO-Spatial \
--seeds 42 123 2024 \
--output outputs/eval/libero_spatial/
# D.3 SimplerEnv (§6.1 universality)
bash scripts/eval/run_simpler.sh \
--checkpoint checkpoints/pace_v2_libero_long \
--output outputs/eval/simpler/
# D.4 LIBERO-Perturbed (robustness)
python scripts/eval/libero_perturbed.py \
--checkpoint checkpoints/pace_v2_libero_long \
--output outputs/eval/libero_perturbed/After Part C produces a checkpoint and Part D produces evaluation JSONs, the ablation matrix and the phenomenon experiments fold both into the paper figures.
# E.1 Seven-config ablation matrix (21 GPU runs total)
for config in configs/ablation/v2/0[1-7]_*.yaml; do
for seed in 42 123 2024; do
python scripts/training/train_dummy_batch.py \
--config $config --seed $seed --device cuda \
--total_steps 20000 \
--output outputs/ablation_v2/$(basename $config .yaml)/seed_$seed/
done
done
# E.2 Aggregate ablation results (rliable IQM + bootstrap CI + Wilcoxon)
python scripts/aggregate_ablation.py \
--input_root outputs/ablation_v2/ \
--output paper_figures/ablation_v2/
# Produces: ablation_stats.csv, ablation_table_v2.tex (paper Table 2)
# E.3 Phenomenon experiments (§6 of paper)
python scripts/phenomenon/universality.py \
--n_rollouts 50 --seeds 0 1 2 \
--output paper_figures/universality/
python scripts/phenomenon/regret_scaling.py \
--checkpoint checkpoints/pace_v2_libero_long \
--H_values 4 8 16 32 64 \
--output paper_figures/regret_scaling/
python scripts/phenomenon/triangulation_concordance.py \
--checkpoint checkpoints/pace_v2_libero_long \
--output paper_figures/triangulation/
# E.4 Diagnostic measurements (§6.4, §6.5, inference cost)
python scripts/diagnostics/diagnostic_utils/trigger_comparison.py \
--checkpoint checkpoints/pace_v2_libero_long \
--output paper_figures/diagnostics/
python scripts/diagnostics/diagnostic_utils/measure_boundary_error.py \
--checkpoint checkpoints/pace_v2_libero_long \
--output paper_figures/diagnostics/boundary_ratio.json
python scripts/diagnostics/diagnostic_utils/measure_inference_cost.py \
--checkpoint checkpoints/pace_v2_libero_long \
--device cuda --output paper_figures/diagnostics/inference_cost.json
# E.5 Generate paper figures (all read from paper_figures/*/ produced above)
python scripts/figures/fig1_universality.py \
--input paper_figures/universality/raw_distances.json \
--output paper_figures/fig1_universality.pdf
python scripts/figures/fig2_method_overview.py # purely diagrammatic
python scripts/figures/fig3_phase_visualization.py \
--input paper_figures/triangulation/phase_vis_data.json \
--output paper_figures/fig3_phase_visualization.pdf
python scripts/figures/fig4_regret_scaling.py \
--input paper_figures/regret_scaling/regret_vs_H.csv \
--output paper_figures/fig4_regret_scaling.pdf
python scripts/figures/fig5_concordance_pr_curve.py \
--input paper_figures/diagnostics/trigger_comparison.csv \
--output paper_figures/fig5_concordance_pr_curve.pdf
# E.6 One-shot orchestrator (does C.E1 → E.5 with default settings)
bash scripts/run_experiments.sh --checkpoint checkpoints/pace_v2_libero_longOnce Part D and Part E complete, the README placeholders are replaced as follows:
| Placeholder | Replace from |
|---|---|
| Table 1 (LIBERO-Long SR) | outputs/eval/libero_long/eval_results.json (config 07) and outputs/eval/libero_long/*config_06* (oracle) |
| Table 2 (ablation) | paper_figures/ablation_v2/ablation_stats.csv |
| §6.1 universality KS test | paper_figures/universality/raw_distances.json |
| §6.2 regret table | paper_figures/regret_scaling/regret_vs_H.csv |
| §6.3 triangulation P/R/F1 | paper_figures/triangulation/concordance_pr.json |
| §6.4 detector comparison | paper_figures/diagnostics/trigger_comparison.csv |
| §6.5 boundary ratio | paper_figures/diagnostics/boundary_ratio.json |
| Inference cost latency/Hz | paper_figures/diagnostics/inference_cost.json |
A long-horizon manipulation policy executing action chunks faces a fundamental challenge: the flow-matching action distribution shifts sharply at phase boundaries (grasp → transport → place), but the policy must commit to a chunk of H actions before it can detect the shift. We call these boundaries Predictability Cliffs.
PACE v2 addresses this with three components:
-
Hierarchical Phase Encoder — a two-level FSQ (macro K₁ = 20, micro K₂ = 30) that makes phase structure identifiable across seeds via InfoNCE.
-
Three Cliff Estimators and Concordance — I^(1) (Bhattacharyya β_t), I^(2) (action-ensemble variance), I^(3) (velocity-field curvature) are fused via rank-based concordance C_t. A cliff is detected when all three estimators agree.
-
PACE Closed-loop Replanning (PCAR) — adaptive replanning triggered by C_t, with a budget constraint to prevent over-triggering.
Boundary-aware flow loss reweighting (w(β) = 1 + λ·β) further improves action prediction quality near phase transitions.
For the full mathematical derivation and ablation design, see docs/ARCHITECTURE.md.
configs/
train/ Three-stage training curriculum (01–03)
ablation/v2/ Seven ablation configs + 2 sanity checks
baselines/ Cross-policy adapters (OpenVLA, π0, BC-ACT, Diffusion Policy)
lerobot_policy_phaseqflow/
src/.../ Policy implementation (PhaseQFlowPolicy, PhaseQFlowConfig)
phase_centric/ Training-side: phase posterior, cliff estimators (loss path),
PCAR / B-PCAR / boundary-aware flow loss
inference/ Runtime-side (PACE v2): compute_policy_variance (I^(2)),
compute_velocity_curvature (I^(3)), ConcordanceDetector (C_t)
scripts/
phenomenon/ Universality, regret scaling, triangulation experiments
eval/ LIBERO-Perturbed and SimplerEnv evaluation
figures/ Publication-grade figure generation
diagnostics/ Boundary error, replan alignment, trigger comparison, cost
calibration/ Concordance threshold and B-PCAR sweeps
tests/ Unit + smoke tests (covers training and inference paths)
docs/
ARCHITECTURE.md Full architecture specification
OPERATIONS_GUIDE.md Engineering handbook
EXPERIMENT_RUNBOOK.md End-to-end GPU reproduction guide
The training and inference cliff signals are intentionally split:
phase_centric/cliff_estimators.pyprovides tensor-level functions (compute_I_hat_1/2/3,compute_concordance_C) consumed insidePhaseQFlowPolicy.forward()for loss / diagnostics.inference/exposes the runtime API used by evaluation rollouts (scripts/eval/libero_perturbed.py) —ConcordanceDetector+ standalone estimator functions that read the policy's cached_last_betaandflow_action_head._last_conditionto fuse I^(1)/I^(2)/I^(3) into C_t.
@inproceedings{pace2027,
title = {{PACE}: Predictability-Aware Closed-loop Execution
for Long-Horizon Manipulation},
author = {[Authors]},
booktitle = {Conference on Robot Learning (CoRL)},
year = {2027},
}