Repository navigation
Archives and Replay
aporia-store preserves benchmark runs so their outputs can be checked against the recorded A-IR. A run archive contains a manifest with model and experiment context, SHA-256 digests, fixed-record observations, findings, and Atlas bands in CSV form. The records include bit-level numerical data rather than relying only on rounded report text.
Replay loads the stored A-IR and observation records, checks the archive's manifest digests, then re-executes the recorded inputs. It compares outputs and traces by floating-point bit pattern and also checks raised flags and instruction-step counts. A mismatch is reported as a replay mismatch, not silently rounded away. The archive's replay path is exposed by the store library and used by the benchmark harness.
Replay is intentionally scoped. It checks whether the stored model and execution path reproduce the recorded observations under the replay configuration. It does not establish that the model is physically correct, that the declared assumptions are true, or that a different external program has been captured in the archive.
The benchmark harness archives each run at the largest budget for seed 1, validates the archive, and replays it before reporting. In the current results, 63 of 63 archived runs replayed bit-for-bit across 37,506 executions. This is evidence about the committed run records, not a claim about all future runs or platforms. See Experiments & Results.
The current .ap CLI does not write an archive. The benchmarking workflow does. The store can produce a separate replayable findings representation, and benchmark-generated findings may include counterexample reductions from Counterexample Minimisation.
Home · Architecture · Counterexample Minimisation · Benchmarking