This repository is the artifact for the paper:
SVN: Shape Value Numbering for Comprehensive and Practical Safety Assessment
Submitted to CGO 2027
It contains the benchmark suite, evaluation scripts, and build orchestration to reproduce the results (RQ1–RQ4) presented in the paper.
.
├── choreo/ # Choreo compiler (git submodule → GitHub)
├── benchmark/
│ ├── choreo/ # 310 Choreo (.co) benchmark cases (15 categories)
│ ├── mlir/ # MLIR linalg comparison cases
│ ├── memref/ # MLIR memref comparison cases
│ ├── iree/ # IREE comparison cases
│ ├── triton/ # Triton comparison cases
│ ├── bugs/ # Bug injection mutants (RQ2)
│ └── results/ # (generated locally, not committed)
├── scripts/ # Data collection, plotting, and automation
│ ├── reproduce_all.sh # ★ One-command reproduction script
│ ├── choreo_assertion_stats.py # RQ1: assessment coverage & discharge
│ ├── bug_detection_eval.py # RQ2: bug detection effectiveness
│ ├── choreo_compile_overhead.py # RQ4: compile-time overhead
│ ├── choreo_runtime_entry.py # RQ3: runtime assertion overhead
│ ├── visualize_results.py # Terminal + HTML report generation
│ ├── collect_all_stats.py # Cross-system comparison
│ └── ...
├── Makefile # Build targets
└── README.md # This file
| Tool | Version | Notes |
|---|---|---|
| GCC / G++ | >= 9.0 | C++17 support required |
| CMake | >= 3.16 | Build system |
| Ninja | any | ninja-build package |
| Python | >= 3.8 | For statistics and plotting scripts |
| matplotlib | any | Optional: for PNG figures and HTML report |
| Git | any | Submodule checkout |
| flex/bison | >= 2.6/3.8 | Auto-downloaded if missing (see below) |
| CUDA | >= 12.0 | Required for RQ3 (runtime overhead) + GPU tests |
Flex and Bison are auto-downloaded and compiled from source during CMake configuration if the system versions are missing or too old.
git clone --recursive https://github.com/LancerLab/svn-artifacts.git
cd svn-artifacts
bash scripts/reproduce_all.shThis will:
- Initialize the Choreo submodule and its dependencies (cutlass, gtest)
- Build Choreo from source
- Run compile-time tests (check + cli)
- Collect RQ1 assessment statistics (310 cases × 15 categories)
- Run RQ2 bug detection evaluation (210 injected bugs × 3 systems)
- Measure RQ3 runtime assertion overhead (if CUDA GPU available)
- Measure RQ4 compile-time overhead (153 symbolic cases)
- Print a comparison table against the paper values
- Generate an interactive HTML report (
benchmark/results/report.html)
Results are written to benchmark/results/.
| Step | Approx. Time | Notes |
|---|---|---|
| Build Choreo | ~1 min | Parallel make |
| Compile-time tests | ~30 sec | lit runner |
| RQ1: Assessment stats | ~2 min | 310 cases, parallel |
| RQ2: Bug detection | ~5 min | 210 mutants × SVN + MLIR |
| RQ3: Runtime overhead | ~10 min | GPU required, 7 reps/case |
| RQ4: Compile overhead | ~3 min | 153 cases × 5 reps |
| MLIR build (optional) | ~30 min | Full LLVM from source |
| Visualization | ~10 sec | Report generation |
| Total (with GPU) | ~22 min | Excluding optional MLIR build |
The script produces:
- Terminal: Rich summary tables with per-category breakdowns for all RQs
benchmark/results/report.html: Self-contained HTML with interactive Chart.js graphsbenchmark/results/figures/: PNG figures for each RQ (requires matplotlib)benchmark/results/choreo_stats.csv: Raw RQ1 databenchmark/results/bug_detection_results.csv: Raw RQ2 databenchmark/results/choreo_runtime_entry.csv: Raw RQ3 data (if GPU available)benchmark/results/choreo_compile_overhead.csv: Raw RQ4 databenchmark/results/reproduce_all.log: Full terminal log of the reproduction run
After reproduction completes, you can re-display the full summary at any time without re-running experiments:
python3 scripts/show_results.pyThis reads the CSV files in benchmark/results/ and prints the same detailed
tables (RQ1–RQ4 breakdowns, paper-vs-reproduced comparison). Use --no-color
for pipe-friendly output, or --results-dir DIR to point at a different data
directory.
Evaluates the breadth (ACD: Assessment Coverage Density) and resolution capability (ADR: Assessment Discharge Ratio) of each system.
| Metric | Paper Value |
|---|---|
| Total assessments | 12,592 |
| Static discharged | 11,753 |
| ADR | 93.3% |
| Cases compiled | 310/310 |
| ACD | 40.6/case |
Comparison: MLIR generates 2,634 (ACD 8.5, ADR 62.9%), IREE generates 370 (ACD 1.2, ADR 0%), Triton generates 0 compiler assessments (manual only).
Tests detection of 210 injected shape bugs across 4 classes:
- Dimension mismatch (139 bugs)
- Input-dependent OOB (58 bugs)
- Wrong output shape (8 bugs)
- Stride/layout error (5 bugs)
| System | Detected | BDE | Resolution |
|---|---|---|---|
| SVN | 210/210 | 100% | All compile-time |
| MLIR | 139/210 | 66.2% | 80 static + 59 runtime |
| IREE | 80/210 | 38.1% | All entry-level |
Measures execution-time overhead (RAO) at four assertion levels:
- none: baseline (no assertions)
- entry: host-side entry-point checks only
- all (hoisted): full checks with assertion hoisting
- all (no-hoist): full checks without hoisting
| Level | Paper Avg | Paper Max |
|---|---|---|
| Entry | <0.4% | — |
| All (hoisted) | +1.8% | +7.1% |
| All (no-hoist) | +9.6% | +92.6% |
Hoisting delivers a 5.3x cost reduction.
Measures SVN's frontend compilation cost on 153 symbolic-dimension cases.
| Metric | Paper Value |
|---|---|
| CTO | 4.7% |
| Per-case | ~3.7 ms |
# 1. Build Choreo
make choreo-build
# 2. Run compile-time tests
make choreo-test
# 3. Collect assessment statistics (RQ1)
make choreo-stats
# 4. Run bug detection evaluation (RQ2)
python3 scripts/bug_detection_eval.py
# 5. (Requires CUDA GPU) Measure runtime overhead (RQ3)
export CUDA_HOME=/usr/local/cuda
export CUTE_HOME=$(pwd)/choreo/extern/cutlass
python3 scripts/choreo_runtime_entry.py --reps 7 --levels none,entry,all,all-nohoist
# 6. Measure compile-time overhead (RQ4)
make choreo-cto
# 7. Generate visualization
python3 scripts/visualize_results.py
# 8. (Optional) Cross-system comparison
make mlir-clone && make mlir-build
python3 scripts/collect_all_stats.pyThe cross-system comparison (SVN vs MLIR vs IREE vs Triton) requires building the MLIR tools:
make mlir-clone # shallow-clone llvm-project release/22.x
make mlir-build # build mlir-opt, mlir-translate, FileCheck (~30 min)Then re-run bash scripts/reproduce_all.sh without --skip-mlir.
| Component | Version | Source |
|---|---|---|
| Choreo (SVN) | cgo2027-eval | github.com/LancerLab/croqtile |
| LLVM/MLIR | release/22.x | github.com/llvm/llvm-project |
| IREE | v3.10.0 | pre-compiled or scripts/fetch_mlir_baselines.sh |
| Triton | v3.6.0 | scripts/fetch_mlir_baselines.sh |
| CUTLASS | v4.2.1 | via Choreo submodule |
| GoogleTest | latest | via Choreo submodule |
See individual component licenses. The benchmark cases and evaluation scripts in this repository are provided for artifact evaluation purposes.
The paper numbers were collected on a specific hardware configuration. Reproduced numbers may vary slightly:
-
Assessment count: The current compiler may generate more assessments than the paper reports (the compiler has been improved since paper submission). The ADR (discharge ratio) should remain within 92-94%.
-
CTO: Compile-time overhead is sensitive to machine load, CPU cache state, and measurement repetitions. With 3 reps, variance can mask the small (4.7%) overhead. Use --rq4-reps 10 for more stable measurements.
-
RAO: Entry-level runtime overhead is consistently below 0.4% (typically below 0.1% median). Small negative values indicate measurement noise.
-
Compile failures: Some cases (conv2d, matmul) may fail on machines without certain CUDA capabilities. The -fc (fast-compile) flag maximizes compatibility. Expect at least 305/310 cases to succeed.