Skip to content
bhogesararam23 edited this page Oct 5, 2026 · 1 revision

FAQ

Does TRUSTED mean the model is correct?

No. It means the current evidence model found no problem in that cell under the tested assumptions, samples, enabled channels, and label policy. An unsampled or weakly measured cell may remain UNKNOWN. See Trust Atlas.

What does a strength of 0.8 mean?

It is a calibrated signal strength relative to typical measurements in the experiment. It is not an 80% probability of failure. Fused risk is also a ranking quantity, not a probability. See Evidence Model.

Can I run APORIA on a program that is not written in its DSL?

The campaign library has an Executor interface that can receive results from another execution path. The shipped aporia run command reads .ap files, and the repository does not yet provide a packaged adapter or standalone foreign-program workflow. See Architecture and Limitations.

Does every analysis run produce an archive and a minimised counterexample?

No. aporia run <model.ap> prints a campaign report but does not write an archive or minimise findings. The benchmark harness writes run archives and uses minimisation in its workflow. See Archives and Replay and Counterexample Minimisation.

Does a compile-time error count as a detected runtime fault?

No. The benchmark verifies the model's declared ground truth before running campaigns. The unit-mistake corpus entry is rejected by the unit checker and is excluded from the swept result matrix.

Is adaptive search better than random or stratified search?

Not generally based on the committed measurements. In the current comparison, adaptive is better on one entry, worse on two, and equal or unresolved on fifteen. It localises the narrow RLC resonance case for one seed at the largest budget, while coverage strategies do better on other cases. See Experiments & Results.

Why does an obviously odd region sometimes stay UNKNOWN?

One channel can report an unusual behavior that is real but not a model defect. By default, measurement-based suspicion needs two channels speaking about the same evaluation, while cells also need enough samples before they can be trusted. That policy accepts more unknown area to avoid treating an isolated signal as a defect. See Evidence Model.

Are benchmark times comparable across runs?

The source results document a substantial wall-time change between otherwise comparable runs without establishing the cause. Use instruction_steps for deterministic work comparisons within the recorded ladder, and treat elapsed time as machine- and run-dependent. See Benchmarking.

What does replay prove?

Replay checks the stored model and inputs against recorded outputs, traces, flags, and step counts, including bit patterns. It establishes reproducibility of those recorded executions under the replay setup, not correctness of the scientific assumptions or outputs. See Archives and Replay.


Home · Architecture · Limitations · Roadmap

Clone this wiki locally