Alice 0.2.6
Alice 0.2.6 — Continuable Cooling, Faster Thermal Measurement
Release Date: August 20, 2026
Version 0.2.6 generalizes XTRG resumption: run() now accepts a starting state from
any archived artifacts/step_XX.ckpt, not just the step the previous run stopped at.
A finished run can therefore be cooled further, or a segment re-cooled under different
options, with the β/log Z history in thermal.ckpt truncated and recomputed as needed.
Separately, observe() on a thermal NormalMPO is rewritten as a transfer-matrix
environment sweep instead of forming and compressing the MPO product ρ · O. No
breaking API changes.
🔁 Continuing a Finished XTRG Run
run()accepts a state at any step covered bythermal.ckpt, including one before
the end of that history. The entries paststate.stepare truncated (with a
WARNINGnaming how many are dropped), then recomputed by this run and overwritten
on disk together with theirstep_XX.ckptarchives. Previously the history had to
end exactly atstate.step, so only the interrupted-run case worked.- History recovery moves out of
run()into a new_resume_historyhelper, which
validates the recoveredSummaryand returns thebetas/log_z/discarded_weights
prefixes for steps0 … state.step. - Validation is now anchored at
state.steprather than at the end of the history: β is
compared athistory.betas[state.step]instead ofhistory.betas[-1], and the
history is rejected only if it stops beforestate.step. - New τ₀ consistency check:
history.betas[0]must matchopts.tau_0, so continuing a
run with options built on a different τ₀ fails immediately instead of silently
producing a summary whose β grid does not correspond to its owntau_0. - New guard on
state.step > opts.n_steps, raisingValueErrorwith a reminder that
n_stepsis the absolute step index to stop at, counted from τ₀ — not a number of
additional steps to perform.state.step == opts.n_stepsremains a valid no-op. - The startup banner on a resumed run reports the starting step and the number of steps
remaining, and derivesbeta_maxfromstate.betarather than from τ₀, which
describes the full run only when starting at step 0.
🌡️ Thermal observe() via Environment Sweep
observe(rho, O)for aNormalMPOnow evaluatesTr[ρ O] / Tr[ρ]with a
left-to-right sweep accumulating a second-order environmentE[ρ_bond, O_bond],
mirroring the existing MPS–MPO–MPS path. A single per-siteeinsumcontracts ρ's
phys_outagainst O'sphys_in(the matrix product) and ρ'sphys_inagainst O's
phys_out(closing the trace loop).- This replaces forming
ρ · Oand callingcompact()on it, which built an MPO of
bond dimensionχ_ρ · χ_Oand paid a compression sweep before the trace — work the
environment sweep never materializes. - Numerator magnitude is recovered as
log|raw| + rho.log_scale + O_norm.log_scale, and
the ratio is still combined in log-space so it stays correct when either trace alone
would overflow float64 — the normal situation deep into an XTRG cooling run. - The right-boundary scalar extraction accounts for the Bridge (
intw) normalization
weight, so generic (non-Abelian) symmetry groups are handled alongside Abelian ones. - A length mismatch between
rhoandobservablenow raises aValueErrornaming both
lengths, instead of failing further inside the contraction.
🧪 Tests
- New
TestResumecases intest_xtrg.py:
test_continue_finished_run_from_archived_artifact(continue a finished run from
step_02.ckptat a largermax_bond, checking the earlier grid points survive
untouched),test_continue_from_earlier_step_truncates_history(re-cool steps 3–4 and
assert the truncation warning), plustest_history_shorter_than_state_step_raises,
test_mismatched_tau_0_raises, andtest_state_step_past_n_steps_raises. - New thermal
observe()coverage intest_observe.py, written once in a shared
_ObserveThermalTestsbase and run against both U(1) and SU(2) realizations of the
same Heisenberg chain: finite-float return, length-mismatch and zero-norm errors, and
linearity in the observable. TestObserveThermalSU2additionally cross-checks⟨H⟩_βbetween the SU(2) and U(1)
builds of the same physical Hamiltonian — the strongest correctness signal available
without a closed-form reference, and the only test exercising the Bridge-weight branch.- New session-scoped
heisenberg_mpo_u1/heisenberg_mpo_su2fixtures in
tests/network/conftest.py, built from anL = 4nearest-neighbor chain.
📚 Documentation
xtrg/index.mdgains a "Continuing a finished run" section: loading an archived
step_XX.ckpt, the absolute meaning ofn_steps, the τ₀ requirement, and re-cooling
a segment from an earlier step.- The
xtrgmodule docstring andrun()'s docstring document the same, including the
newValueErrorconditions. xtrg/summary.mdnotes that a continued run appends to the recovered history, so one
summary may merge segments computed under different options, with no marker at the
junction andu/c_Vfinite differences there mixing both accuracies — copy the
checkpoint directory first to keep the original series.
📊 Statistics
- 968 tests across 29 test modules (up from 954 / 29 modules in v0.2.5).
- 8 commits since v0.2.5.
- 8 files changed, 498 insertions, 46 deletions.
- 28 source modules in four subpackages:
alice.network,alice.physics,
alice.algorithm.dmrg,alice.algorithm.xtrg(unchanged from v0.2.5).
✅ Compatibility
Breaking Changes: none.
Behavioral Changes:
- XTRG resumption is more permissive in one direction and stricter in two others: a
thermal.ckptreaching paststate.stepis now accepted (and truncated) rather than
rejected, while a τ₀ disagreeing with the history and astate.steppast
opts.n_stepsare now rejected. Existing interrupted-run resumption fromxtrg.ckpt
is unaffected. - A continued run overwrites the
step_XX.ckptarchives and thethermal.ckptentries
past its starting step. Copy the checkpoint directory first if the earlier series is
worth keeping. - Thermal
observe()results may differ from v0.2.5 in the last few digits, since the
environment sweep skips thecompact()compression the MPO-product path applied
before tracing.
Requirements:
- Python ≥ 3.11
- PyTorch ≥ 2.5
- Nicole ≥ 0.3.7
📝 Notes
Both changes are aimed at long XTRG runs. The resumption logic introduced in v0.2.4
solved a narrow problem — a crashed run picking up where it left off — but assumed the
history on disk ended exactly at the state being handed back, which made the far more
common workflows impossible: a run that reached β = 2²⁰τ₀ and now needs to go to 2²⁵τ₀,
or one whose last few steps were clearly under-converged and should be redone at a
larger bond dimension. Anchoring validation at state.step and truncating rather than
refusing turns the archived step_XX.ckpt files into genuine restart points, which is
what they were always meant to be. The observe() rewrite matters for the same runs
from the other direction: measuring an observable on a cold ρ with bond dimension in
the hundreds should not cost a full MPO product plus compression sweep, and now it
costs one environment sweep instead.