EuroSys 2027 note: for the EuroSys 2027 paper (The Compiler as Auditor: Safety Ledgers for AI-Native GPU Kernels), use
bash scripts/reproduce_asplos27.sh # add --skip-build / --skip-gpu as neededwhich regenerates every CSV behind the paper's tables and figures (RQ1–RQ6), re-renders them into
../asplos27/{tables,figures}/, and prints a number-by-number comparison against the paper.⚠ Known gap. This script still renders into
../asplos27/. The paper's active tree is now../eurosys27/, so areproduce_eurosys27.sh(or a retarget of this one) is required before submission — acceptance condition C3 in../eurosys27/plan/acceptable-e2e-test.md. Until then, reproduction writes into a superseded tree.The CGO 2027 pipeline (
reproduce_all.sh) is retained for reference; its summary table compares against the older CGO numbers.Evaluation status. The paper's numbers are currently gated by
../eurosys27/plan/data-integrity-blockers.md(D0–D7) and../eurosys27/plan/acceptable-e2e-test.md. Mutation-generation coverage is being expanded underbenchmark2/specs/mutation-specs-v2.md; the frozenbenchmark2/specs/mutation-specs.md(v1.0) remains the spec that all existingstats.jsonwere produced under.
This repository is the artifact for the paper:
SVN: Shape Value Numbering for Comprehensive and Practical Safety Assessment
Submitted to CGO 2027 (rejected)
and, under a reworked framing, for its successor:
The Compiler as Auditor: Safety Ledgers for AI-Native GPU Kernels
Target: EuroSys 2027
It contains the benchmark suite, evaluation scripts, and build orchestration to reproduce the results (RQ1–RQ4) presented in the paper.
.
├── croqtile/ # Croqtile compiler (git submodule → GitHub)
├── benchmark/
│ ├── choreo/ # 310 Choreo (.co) benchmark cases (15 categories)
│ ├── mlir/ # MLIR linalg comparison cases
│ ├── memref/ # MLIR memref comparison cases
│ ├── iree/ # IREE comparison cases
│ ├── triton/ # Triton comparison cases
│ ├── bugs/ # Bug injection mutants (RQ2)
│ └── results/ # (generated locally, not committed)
├── scripts/ # Data collection, plotting, and automation
│ ├── reproduce_all.sh # ★ One-command reproduction script
│ ├── choreo_assertion_stats.py # RQ1: assessment coverage & discharge
│ ├── bug_detection_eval.py # RQ2: bug detection effectiveness
│ ├── choreo_compile_overhead.py # RQ4: compile-time overhead
│ ├── choreo_runtime_entry.py # RQ3: runtime assertion overhead
│ ├── visualize_results.py # Terminal + HTML report generation
│ ├── collect_all_stats.py # Cross-system comparison
│ └── ...
├── Makefile # Build targets
└── README.md # This file
| Tool | Version | Notes |
|---|---|---|
| GCC / G++ | >= 9.0 | C++17 support required |
| CMake | >= 3.16 | Build system |
| Ninja | any | ninja-build package |
| Python | >= 3.8 | For statistics and plotting scripts |
| matplotlib | any | Optional: for PNG figures and HTML report |
| Git | any | Submodule checkout |
| flex/bison | >= 2.6/3.8 | Auto-downloaded if missing (see below) |
| CUDA | >= 12.0 | Required for RQ3 (runtime overhead) + GPU tests |
Flex and Bison are auto-downloaded and compiled from source during CMake configuration if the system versions are missing or too old.
git clone --recursive https://github.com/LancerLab/svn-artifacts.git
cd svn-artifacts
bash scripts/reproduce_all.shThis will:
- Initialize the Choreo submodule and its dependencies (cutlass, gtest)
- Build Choreo from source
- Run compile-time tests (check + cli)
- Collect RQ1 assessment statistics (310 cases × 15 categories)
- Run RQ2 bug detection evaluation (210 injected bugs × 3 systems)
- Measure RQ3 runtime assertion overhead (if CUDA GPU available)
- Measure RQ4 compile-time overhead (153 symbolic cases)
- Print a comparison table against the paper values
- Generate an interactive HTML report (
benchmark/results/report.html)
Results are written to benchmark/results/.
| Step | Approx. Time | Notes |
|---|---|---|
| Build Choreo | ~1 min | Parallel make |
| Compile-time tests | ~30 sec | lit runner |
| RQ1: Assessment stats | ~2 min | 310 cases, parallel |
| RQ2: Bug detection | ~5 min | 210 mutants × SVN + MLIR |
| RQ3: Runtime overhead | ~10 min | GPU required, 7 reps/case |
| RQ4: Compile overhead | ~3 min | 153 cases × 5 reps |
| MLIR build (optional) | ~30 min | Full LLVM from source |
| Visualization | ~10 sec | Report generation |
| Total (with GPU) | ~22 min | Excluding optional MLIR build |
The script produces:
- Terminal: Rich summary tables with per-category breakdowns for all RQs
benchmark/results/report.html: Self-contained HTML with interactive Chart.js graphsbenchmark/results/figures/: PNG figures for each RQ (requires matplotlib)benchmark/results/choreo_stats.csv: Raw RQ1 databenchmark/results/bug_detection_results.csv: Raw RQ2 databenchmark/results/choreo_runtime_entry.csv: Raw RQ3 data (if GPU available)benchmark/results/choreo_compile_overhead.csv: Raw RQ4 databenchmark/results/reproduce_all.log: Full terminal log of the reproduction run
After reproduction completes, you can re-display the full summary at any time without re-running experiments:
python3 scripts/show_results.pyThis reads the CSV files in benchmark/results/ and prints the same detailed
tables (RQ1–RQ4 breakdowns, paper-vs-reproduced comparison). Use --no-color
for pipe-friendly output, or --results-dir DIR to point at a different data
directory.
Evaluates the breadth (ACD: Assessment Coverage Density) and resolution capability (ADR: Assessment Discharge Ratio) of each system.
| Metric | Paper Value |
|---|---|
| Total assessments | 12,592 |
| Static discharged | 11,753 |
| ADR | 93.3% |
| Cases compiled | 310/310 |
| ACD | 40.6/case |
Comparison: MLIR generates 2,634 (ACD 8.5, ADR 62.9%), IREE generates 370 (ACD 1.2, ADR 0%), Triton generates 0 compiler assessments (manual only).
Tests detection of 210 injected shape bugs across 4 classes:
- Dimension mismatch (139 bugs)
- Input-dependent OOB (58 bugs)
- Wrong output shape (8 bugs)
- Stride/layout error (5 bugs)
| System | Detected | BDE | Resolution |
|---|---|---|---|
| SVN | 210/210 | 100% | All compile-time |
| MLIR | 139/210 | 66.2% | 80 static + 59 runtime |
| IREE | 80/210 | 38.1% | All entry-level |
Measures execution-time overhead (RAO) at four assertion levels:
- none: baseline (no assertions)
- entry: host-side entry-point checks only
- all (hoisted): full checks with assertion hoisting
- all (no-hoist): full checks without hoisting
| Level | Paper Avg | Paper Max |
|---|---|---|
| Entry | <0.4% | — |
| All (hoisted) | +1.8% | +7.1% |
| All (no-hoist) | +9.6% | +92.6% |
Hoisting delivers a 5.3x cost reduction.
Measures SVN's frontend compilation cost on 153 symbolic-dimension cases.
| Metric | Paper Value |
|---|---|
| CTO | 4.7% |
| Per-case | ~3.7 ms |
# 1. Build Choreo
make choreo-build
# 2. Run compile-time tests
make choreo-test
# 3. Collect assessment statistics (RQ1)
make choreo-stats
# 4. Run bug detection evaluation (RQ2)
python3 scripts/bug_detection_eval.py
# 5. (Requires CUDA GPU) Measure runtime overhead (RQ3)
export CUDA_HOME=/usr/local/cuda
export CUTE_HOME=$(pwd)/choreo/extern/cutlass
python3 scripts/choreo_runtime_entry.py --reps 7 --levels none,entry,all,all-nohoist
# 6. Measure compile-time overhead (RQ4)
make choreo-cto
# 7. Generate visualization
python3 scripts/visualize_results.py
# 8. (Optional) Cross-system comparison
make mlir-clone && make mlir-build
python3 scripts/collect_all_stats.pyThe cross-system comparison (SVN vs MLIR vs IREE vs Triton) requires building the MLIR tools:
make mlir-clone # shallow-clone llvm-project release/22.x
make mlir-build # build mlir-opt, mlir-translate, FileCheck (~30 min)Then re-run bash scripts/reproduce_all.sh without --skip-mlir.
| Component | Version | Source |
|---|---|---|
| Choreo (SVN) | cgo2027-eval | github.com/LancerLab/croqtile |
| LLVM/MLIR | release/22.x | github.com/llvm/llvm-project |
| IREE | v3.10.0 | pre-compiled or scripts/fetch_mlir_baselines.sh |
| Triton | v3.6.0 | scripts/fetch_mlir_baselines.sh |
| CUTLASS | v4.2.1 | via Choreo submodule |
| GoogleTest | latest | via Choreo submodule |
See individual component licenses. The benchmark cases and evaluation scripts in this repository are provided for artifact evaluation purposes.
The paper numbers were collected on a specific hardware configuration. Reproduced numbers may vary slightly:
-
Assessment count: The current compiler may generate more assessments than the paper reports (the compiler has been improved since paper submission). The ADR (discharge ratio) should remain within 92-94%.
-
CTO: Compile-time overhead is sensitive to machine load, CPU cache state, and measurement repetitions. With 3 reps, variance can mask the small (4.7%) overhead. Use --rq4-reps 10 for more stable measurements.
-
RAO: Entry-level runtime overhead is consistently below 0.4% (typically below 0.1% median). Small negative values indicate measurement noise.
-
Compile failures: Some cases (conv2d, matmul) may fail on machines without certain CUDA capabilities. The -fc (fast-compile) flag maximizes compatibility. Expect at least 305/310 cases to succeed.