Repository navigation
FAQ
No. It means the current evidence model found no problem in that cell under the tested assumptions, samples, enabled channels, and label policy. An unsampled or weakly measured cell may remain UNKNOWN. See Trust Atlas.
It is a calibrated signal strength relative to typical measurements in the experiment. It is not an 80% probability of failure. Fused risk is also a ranking quantity, not a probability. See Evidence Model.
The campaign library has an Executor interface that can receive results from another execution path. The shipped aporia run command reads .ap files, and the repository does not yet provide a packaged adapter or standalone foreign-program workflow. See Architecture and Limitations.
No. aporia run <model.ap> prints a campaign report but does not write an archive or minimise findings. The benchmark harness writes run archives and uses minimisation in its workflow. See Archives and Replay and Counterexample Minimisation.
No. The benchmark verifies the model's declared ground truth before running campaigns. The unit-mistake corpus entry is rejected by the unit checker and is excluded from the swept result matrix.
Not generally based on the committed measurements. In the current comparison, adaptive is better on one entry, worse on two, and equal or unresolved on fifteen. It localises the narrow RLC resonance case for one seed at the largest budget, while coverage strategies do better on other cases. See Experiments & Results.
One channel can report an unusual behavior that is real but not a model defect. By default, measurement-based suspicion needs two channels speaking about the same evaluation, while cells also need enough samples before they can be trusted. That policy accepts more unknown area to avoid treating an isolated signal as a defect. See Evidence Model.
The source results document a substantial wall-time change between otherwise comparable runs without establishing the cause. Use instruction_steps for deterministic work comparisons within the recorded ladder, and treat elapsed time as machine- and run-dependent. See Benchmarking.
Replay checks the stored model and inputs against recorded outputs, traces, flags, and step counts, including bit patterns. It establishes reproducibility of those recorded executions under the replay setup, not correctness of the scientific assumptions or outputs. See Archives and Replay.
Home · Architecture · Limitations · Roadmap