Repository navigation
Evidence Model
APORIA records evidence as data rather than reducing each execution to a pass/fail flag. An item identifies a channel, a subject (for example a declared constraint, relation, output, or parameter axis), a raw magnitude, a calibrated strength, observation IDs, and an explanation. Provenance is used later to avoid counting the same executions as independent confirmations.
| Channel | Question it asks |
|---|---|
| Behavioral | Does observed behaviour violate a declared relation or show an inferred pattern such as monotonicity, scaling, symmetry, or continuity? |
| Physical | Did a declared constraint fail, or did the computation diverge? |
| Numerical | Does the result change materially between f64 and f32 execution? |
| Differential | Do independent execution paths disagree, such as runtime and reference evaluator? |
| Sensitivity | Does a small parameter change produce an unusually large local output slope? |
The available signals depend on the model and executor. A declared symmetry needs a valid swapped point and spends an additional execution. Numerical evidence is skipped for executors that do not vary with the requested precision. Reference-path comparison is skipped when the executed equations are not represented by the A-IR the reference evaluator would run.
Channels measure different quantities, so raw magnitudes cannot be added directly. For a channel or sufficiently sampled claim, the calibrator uses the median observed positive magnitude as its typical scale, subject to a channel-specific noise floor and a minimum scale guard. A claim with at least eight measurements gets its own typical reference; otherwise it falls back to its channel's reference. Strength is based on excess over typical, transformed to a value in [0, 1]. Fixed facts, such as divergence or a hard rule failure, retain their assigned absolute strength.
This calibration is experiment-relative. A strength of 0.8 does not mean an 80% chance of incorrectness. If everything in a run behaves anomalously in the same way, a relative scale can fail to make that behaviour stand out; the fitted references are preserved so a reader can inspect the basis.
For each channel, fusion keeps its strongest item rather than summing repeated claims. Channels are sorted strongest first. Each later channel is discounted by the greater of its overlap in observation provenance and its estimated non-negative Pearson correlation with already accepted channels, capped at 0.95. Shared provenance uses Jaccard overlap, with a base half-weight except when the observation sets are effectively identical. The discounted contributions are combined using noisy-OR:
risk = 1 - product(1 - accepted_channel_strength)
The score ranks points and cells; it is not a posterior probability. The code also retains per-channel strengths and contribution discounts so a single score does not hide disagreement. The search may use Pareto ranking when preserving that channel-by-channel view is useful.
The default Atlas policy requires either mean risk at or above 0.45 or peak risk at or above 0.7. A point with risk at least 0.45 makes a cell suspicious if it carries an absolute fact or if at least two channels spoke about that same evaluation. The two-channel rule prevents one loud but non-defect signal from making the whole cell suspicious. A cell not classified suspicious can be trusted only after at least four samples and evidence measurement from at least one channel; otherwise it stays unknown. See Trust Atlas.
These thresholds describe current implementation defaults. The benchmark harness records policy and calibration context in its run data; consumers should read the exact run manifest rather than assume all configurations use defaults.