Skip to content

Alice 0.2.6

Choose a tag to compare

@changkai-zhang changkai-zhang released this 20 Aug 13:04
· 25 commits to stable since this release

Alice 0.2.6 — Continuable Cooling, Faster Thermal Measurement

Release Date: August 20, 2026

Version 0.2.6 generalizes XTRG resumption: run() now accepts a starting state from
any archived artifacts/step_XX.ckpt, not just the step the previous run stopped at.
A finished run can therefore be cooled further, or a segment re-cooled under different
options, with the β/log Z history in thermal.ckpt truncated and recomputed as needed.
Separately, observe() on a thermal NormalMPO is rewritten as a transfer-matrix
environment sweep instead of forming and compressing the MPO product ρ · O. No
breaking API changes.

🔁 Continuing a Finished XTRG Run

  • run() accepts a state at any step covered by thermal.ckpt, including one before
    the end of that history. The entries past state.step are truncated (with a
    WARNING naming how many are dropped), then recomputed by this run and overwritten
    on disk together with their step_XX.ckpt archives. Previously the history had to
    end exactly at state.step, so only the interrupted-run case worked.
  • History recovery moves out of run() into a new _resume_history helper, which
    validates the recovered Summary and returns the betas/log_z/discarded_weights
    prefixes for steps 0 … state.step.
  • Validation is now anchored at state.step rather than at the end of the history: β is
    compared at history.betas[state.step] instead of history.betas[-1], and the
    history is rejected only if it stops before state.step.
  • New τ₀ consistency check: history.betas[0] must match opts.tau_0, so continuing a
    run with options built on a different τ₀ fails immediately instead of silently
    producing a summary whose β grid does not correspond to its own tau_0.
  • New guard on state.step > opts.n_steps, raising ValueError with a reminder that
    n_steps is the absolute step index to stop at, counted from τ₀ — not a number of
    additional steps to perform. state.step == opts.n_steps remains a valid no-op.
  • The startup banner on a resumed run reports the starting step and the number of steps
    remaining, and derives beta_max from state.beta rather than from τ₀, which
    describes the full run only when starting at step 0.

🌡️ Thermal observe() via Environment Sweep

  • observe(rho, O) for a NormalMPO now evaluates Tr[ρ O] / Tr[ρ] with a
    left-to-right sweep accumulating a second-order environment E[ρ_bond, O_bond],
    mirroring the existing MPS–MPO–MPS path. A single per-site einsum contracts ρ's
    phys_out against O's phys_in (the matrix product) and ρ's phys_in against O's
    phys_out (closing the trace loop).
  • This replaces forming ρ · O and calling compact() on it, which built an MPO of
    bond dimension χ_ρ · χ_O and paid a compression sweep before the trace — work the
    environment sweep never materializes.
  • Numerator magnitude is recovered as log|raw| + rho.log_scale + O_norm.log_scale, and
    the ratio is still combined in log-space so it stays correct when either trace alone
    would overflow float64 — the normal situation deep into an XTRG cooling run.
  • The right-boundary scalar extraction accounts for the Bridge (intw) normalization
    weight, so generic (non-Abelian) symmetry groups are handled alongside Abelian ones.
  • A length mismatch between rho and observable now raises a ValueError naming both
    lengths, instead of failing further inside the contraction.

🧪 Tests

  • New TestResume cases in test_xtrg.py:
    test_continue_finished_run_from_archived_artifact (continue a finished run from
    step_02.ckpt at a larger max_bond, checking the earlier grid points survive
    untouched), test_continue_from_earlier_step_truncates_history (re-cool steps 3–4 and
    assert the truncation warning), plus test_history_shorter_than_state_step_raises,
    test_mismatched_tau_0_raises, and test_state_step_past_n_steps_raises.
  • New thermal observe() coverage in test_observe.py, written once in a shared
    _ObserveThermalTests base and run against both U(1) and SU(2) realizations of the
    same Heisenberg chain: finite-float return, length-mismatch and zero-norm errors, and
    linearity in the observable.
  • TestObserveThermalSU2 additionally cross-checks ⟨H⟩_β between the SU(2) and U(1)
    builds of the same physical Hamiltonian — the strongest correctness signal available
    without a closed-form reference, and the only test exercising the Bridge-weight branch.
  • New session-scoped heisenberg_mpo_u1/heisenberg_mpo_su2 fixtures in
    tests/network/conftest.py, built from an L = 4 nearest-neighbor chain.

📚 Documentation

  • xtrg/index.md gains a "Continuing a finished run" section: loading an archived
    step_XX.ckpt, the absolute meaning of n_steps, the τ₀ requirement, and re-cooling
    a segment from an earlier step.
  • The xtrg module docstring and run()'s docstring document the same, including the
    new ValueError conditions.
  • xtrg/summary.md notes that a continued run appends to the recovered history, so one
    summary may merge segments computed under different options, with no marker at the
    junction and u/c_V finite differences there mixing both accuracies — copy the
    checkpoint directory first to keep the original series.

📊 Statistics

  • 968 tests across 29 test modules (up from 954 / 29 modules in v0.2.5).
  • 8 commits since v0.2.5.
  • 8 files changed, 498 insertions, 46 deletions.
  • 28 source modules in four subpackages: alice.network, alice.physics,
    alice.algorithm.dmrg, alice.algorithm.xtrg (unchanged from v0.2.5).

✅ Compatibility

Breaking Changes: none.

Behavioral Changes:

  • XTRG resumption is more permissive in one direction and stricter in two others: a
    thermal.ckpt reaching past state.step is now accepted (and truncated) rather than
    rejected, while a τ₀ disagreeing with the history and a state.step past
    opts.n_steps are now rejected. Existing interrupted-run resumption from xtrg.ckpt
    is unaffected.
  • A continued run overwrites the step_XX.ckpt archives and the thermal.ckpt entries
    past its starting step. Copy the checkpoint directory first if the earlier series is
    worth keeping.
  • Thermal observe() results may differ from v0.2.5 in the last few digits, since the
    environment sweep skips the compact() compression the MPO-product path applied
    before tracing.

Requirements:

  • Python ≥ 3.11
  • PyTorch ≥ 2.5
  • Nicole ≥ 0.3.7

📝 Notes

Both changes are aimed at long XTRG runs. The resumption logic introduced in v0.2.4
solved a narrow problem — a crashed run picking up where it left off — but assumed the
history on disk ended exactly at the state being handed back, which made the far more
common workflows impossible: a run that reached β = 2²⁰τ₀ and now needs to go to 2²⁵τ₀,
or one whose last few steps were clearly under-converged and should be redone at a
larger bond dimension. Anchoring validation at state.step and truncating rather than
refusing turns the archived step_XX.ckpt files into genuine restart points, which is
what they were always meant to be. The observe() rewrite matters for the same runs
from the other direction: measuring an observable on a cold ρ with bond dimension in
the hundreds should not cost a full MPO product plus compression sweep, and now it
costs one environment sweep instead.