Repository navigation
Benchmarking
APORIA's benchmark corpus is a set of small, project-authored models in the Aporia DSL. Each entry has a model.ap and a truth.json that declares the region of interest, boundaries where applicable, fault class, and derivation. A registry lists benchmark families and coverage.
Declared regions are axis-aligned boxes in parameter coordinates. Before measurements run, aporia-bench verify evaluates each model on a dense grid and checks that the regions agree with direct evaluation of the model's own rules and divergence. A run refuses to produce measurements if a declared region does not verify. Compile-time errors are not counted as runtime detections; for example, the mutants/unit_mistake entry is rejected by unit checking and is not part of the swept result matrix.
The corpus contains controls with no declared faulty region, wide and narrow regions, and a curved boundary approximated by boxes. The current registry has 22 entries; one compile-time unit-error example is registered but not swept, leaving 21 entries in the measured ladder. The corpus currently covers sign errors, boundary conditions, unit mistakes, cancellation, time-step sensitivity, state-update ordering, overflow, and incorrect domain assumptions. The corpus README identifies incorrect constants, stopping conditions, normalisation, and scalar/SIMD or CPU/GPU divergence as not yet covered.
The committed current ladder uses budgets 40,80,160,320,640, seeds 1,2,3, and strategies adaptive, stratified, and random. Across 21 swept entries this is 945 campaign runs and 189 entry/seed sweeps. The benchmark harness verifies truth before measurement, writes archives, replays the top-budget runs, and reports metrics, comparisons, and explanations.
The repository commands are run from software/:
cargo run --release -p aporia-bench -- list
cargo run --release -p aporia-bench -- verify
cargo run --release -p aporia-bench -- run --budgets 40,80,160,320,640 --seeds 1,2,3
cargo run --release -p aporia-bench -- explain analytic/sqrt_domain --budget 640Detection is intentionally permissive: some suspicious cell must touch a declared region. Localisation is stricter: a suspicious cell must touch a declared region and be at least half inside the declared region. Boundary precision compares produced bands with declared transition values, normalised by axis width and measured only where a band exists. The harness also reports suspicious/trusted/unknown volume, control cleanliness, false-positive metrics, findings, counterexamples, replay, instruction steps, and elapsed time.
Do not read wall-clock time as a stable compute-cost measure. The results README documents a large timing shift between otherwise comparable runs without establishing its cause. instruction_steps is deterministic for the same model, strategy, seed, and budget, and is the preferred work measure in that comparison.
The JSON result files preserve successive measurements and corrections. Only results-1791158752.json is marked current in the results README; its functional fields match the preceding measurement except for counterexamples, while wall_ms is explicitly treated as non-comparable. See Experiments & Results for the interpretation and Limitations for what the suite does not establish.
Home · Adaptive Search · Experiments & Results · Limitations