Skip to content

Experiments and Results

bhogesararam23 edited this page Oct 5, 2026 · 1 revision

Experiments & Results

The results documented here come from the repository's committed benchmark JSON and software/benchmarks/results/README.md. The latest run recorded there is dated 2026-10-05. It measures a specific corpus and build; it does not settle APORIA's general research question.

Current ladder

The current run has 21 swept corpus entries and uses 3 strategies, 3 seeds, and 5 budgets, for 945 campaigns and 189 sweeps. The 22nd registry entry is the unit-mistake example, which the DSL rejects before execution. Across the ladder, 134 sweeps detected a declared region at some budget and 97 localised one.

The strategy comparison reports adaptive better on one entry, worse on two, and equal or unresolved on fifteen. On electromagnetics/rlc_resonance, a band 0.088% of the domain, adaptive localised at budget 640 for one of three seeds; stratified and random detected but did not localise it. Adaptive is not uniformly superior: stratified and random localised synthetic/narrow_1d at 160, while adaptive required 320; stratified localised all three seeds of synthetic/one_pct_3d at 640, while adaptive localised one. One narrow-band win and two baseline wins are not a general result.

Controls and Atlas uncertainty

The corpus has three controls, including electromagnetics/coupled_coils, whose declared symmetry is exercised by swapped-parameter probes. Across the three controls, 118 of 135 campaigns report no suspicious volume; 16 of 27 sweeps are clean at every budget. The remaining flags are small: the notes report about 0.10% mean suspicious volume for control campaigns at budget 640 in the symmetry run.

Corroboration reduces single-channel false alarms but leaves many cells unknown. At budget 640, the reported corpus-wide adaptive coverage is approximately 0.603 trusted and 0.199 unknown. UNKNOWN is a meaningful outcome: it records insufficient or uncorroborated evidence instead of turning a weak single-channel signal into either a defect or a clean bill of health. See Trust Atlas.

Boundary measurement

The current results report 810 boundary-checked rows, 579 with at least one band, and 238 bands within the declared tolerance. Median normalised boundary error is 0.1044 of the axis width. That error is described as coarse and is affected by the axis-aligned partition and minimum-sample floor. The result does not mean that every benchmark boundary was localized.

Replay and minimisation

All 63 of 63 archived top-budget runs replayed bit-for-bit across 37,506 executions. The current run reports 341 of 341 minimisation rows verified with their recorded oracle. Verification means the reduced case preserved that oracle's question; it does not mean every result removed a parameter or is a globally smallest counterexample. See Archives and Replay and Counterexample Minimisation.

What the run changed and what it did not

The results history preserves earlier ladders rather than replacing them. The run that wired declared symmetry added one clean control and showed that the previous 180 sweeps' measured fields were unchanged. The risk-threshold minimisation oracle then changed counterexample results while the 945 campaign outcomes otherwise matched the previous run. The latest Executor refactor was followed by a run: all 945 campaigns matched the preceding campaign fields except counterexamples and wall time. These checks prevent a refactor statement from resting only on an assertion.

The result README also corrects earlier interpretations: control cleanliness improved after two-channel corroboration, a former detection count included noise hits rather than valid localisations, and a previous control table understated suspicious volume. Historical files remain so these changes can be inspected. Wall time varied substantially between same-work runs without a cause established; use instruction steps for deterministic work comparisons. Full per-entry tables and derivations live in the source results README.


Home · Benchmarking · Limitations · Roadmap

Clone this wiki locally