Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
5 changes: 5 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -1 +1,6 @@
resources/DTVM_paper.pdf filter=lfs diff=lfs merge=lfs -text

# Paper reproduction benchmark data (benchmarks/paper)
benchmarks/paper/**/*.wasm filter=lfs diff=lfs merge=lfs -text
benchmarks/paper/**/*.expected filter=lfs diff=lfs merge=lfs -text
benchmarks/paper/**/*.png filter=lfs diff=lfs merge=lfs -text
15 changes: 15 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -47,3 +47,18 @@ bazel-testlogs
.claude/*
!.claude/skills/
!.claude/skills/**

# Paper reproduction (benchmarks/paper) generated artifacts
benchmarks/paper/runtimes/*/build/
benchmarks/paper/runtimes/*/out/
benchmarks/paper/runtimes/dtvm/dtvm
benchmarks/paper/runtimes/dtvm/CMakeCache.txt
benchmarks/paper/runtimes/dtvm_main/dtvm
benchmarks/paper/runtimes/dtvm_main/CMakeCache.txt
benchmarks/paper/raw_data/
benchmarks/paper/scripts/webassembly-testsuites/case/
benchmarks/paper/**/.DS_Store

# benchmarks/paper python artifacts
benchmarks/paper/**/__pycache__/
benchmarks/paper/**/*.pyc
108 changes: 108 additions & 0 deletions benchmarks/paper/ENVIRONMENT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,108 @@
# Experiment Environment and Toolchain

This file collects the environment and version information that can be **confirmed from the archive**. Anything not present in the reproduction notes, the contract-testbed repository, this repository's docs, or the original lab notes is marked **not recorded** rather than guessed.

Different benchmark suites may have used different machines or toolchains; the tables below separate the **contract microbenchmark (vm_benchmark)** from the **other paper baselines**.

---

## Checklist (contract vm_benchmark)

**Test machine (from experiment notes, confirmed):**

> **Compute-optimized c7, 8C16G, Intel(R) Xeon(R) Platinum 8369B CPU @ 2.70GHz**
> Runtime `/proc/cpuinfo`: **cpu MHz = 2699.998**; **Turbo Boost not enabled**

| # | Item | Status | Archived evidence |
|---|------|--------|-------------------|
| — | **Machine hardware** | **confirmed** | see above; source: reproduction notes |
| 1 | Original OS version | **not recorded** | the notes give the hardware spec but no Linux distribution or version |
| 2 | Kernel version | **not recorded** | — |
| 3 | Docker / container | **not recorded** | no container usage recorded |
| 4 | Dockerfile / image / env scripts | **partial** | no Dockerfile; the contract bench ran `sh build.sh release.vmbench 6 false 2` inside the contract-testbed repo (see `benchmarks/contracts/REPRODUCE.md`) |
| 5 | Dedicated machine | **not recorded** | — |
| 6 | CPU core pinning | **not recorded** | the notes say the microbenchmark is "serial, 1C is enough"; no `taskset`/pinning command was recorded |
| 7 | Fixed CPU frequency | **partial** | `/proc/cpuinfo` read **cpu MHz ≈ 2699.998** at runtime (archived runner source, not included); no manual `cpufreq` step was recorded |
| 8 | Turbo Boost disabled | **confirmed** | notes: "turbo boost not enabled"; the archived runner source comments also require disabling Turbo Boost on Intel CPUs |
| 9 | CPU governor | **not recorded** | — |
| 10 | OS / runtime cache clearing | **not recorded** | no `drop_caches` step recorded for the contract bench; **State and Module LRU caches are reused** across measurements (see `benchmarks/contracts/REPRODUCE.md` §5) |
| 11 | Background load | **not recorded** | — |
| 12 | Toolchain versions | **partial** | see "contract toolchain" below; **no** solc / hex-compile command versions |

---

## Contract Microbenchmark (vm_benchmark)

Sources: reproduction notes (2025-04-22 archive) + `benchmarks/contracts/`.

### Hardware and CPU

| Item | Value | Source |
|------|-------|--------|
| Cloud instance | compute-optimized **c7, 8C16G** | experiment notes |
| CPU model | Intel Xeon Platinum **8369B** @ **2.70GHz** (nominal) | experiment notes |
| Observed frequency | **2699.998 MHz** (`cpu MHz` in `/proc/cpuinfo`) | experiment notes + archived runner source |
| Turbo Boost | **off** | experiment notes |
| Workload shape | **serial** microbenchmark; the notes state "**1C is enough**", test environment was **8C16G** | experiment notes |
| OS | **Linux** (reads `/proc/cpuinfo`); distribution **not recorded** | code + gap |

### Repository and DTVM Macros

| Item | Value |
|------|-------|
| Repository | contract-testbed repository |
| Branch | development branch |
| DTVM build macros | `ZEN_ENABLE_JIT`, `ZEN_ENABLE_SINGLEPASS_JIT`, `ZEN_ENABLE_DWASM`, `ZEN_ENABLE_CHECKED_ARITHMETIC` |

### Contract Toolchain

| Tool | Archive status |
|------|----------------|
| **evmone** | pulled in via the testbed's Bazel deps; **no** standalone version pin |
| **DTVM (`dtvm`)** | see the macros above; the `dtvm` binary is built from the DTVM source fork (`VERSIONS.md`) |
| **solc** (hex generation) | **not recorded** |
| **wabt / wat2wasm** | **not used** in the archived contract-bench flow |
| **clang / LLVM** (contract bench itself) | **not recorded** as a standalone version; hex are prebuilt artifacts |

---

## Other Paper Baselines (non-contract)

The following come from experiment notes / `VERSIONS.md` / the experiment notebook, and **may not** share the contract-bench machine.

### Comparison Runtimes

| Runtime | Notes / archived version | Notes |
|---------|--------------------------|-------|
| **Wasmtime** | **31.0.0** ([official x86_64-linux tar.xz](https://github.com/bytecodealliance/wasmtime/releases/download/v31.0.0/wasmtime-v31.0.0-x86_64-linux.tar.xz)) | `runtimes/wasmtime-31.0.0/BUILD.md` |
| **Wasmer** | **`v5.0.4`**; the Singlepass comparison also used **`v5.0.5-rc1`** (wasi limitation) | Wasmer is **not pinned** in this repo |
| **WAMR `iwasm`** | **1.2.3** | `runtimes/wamr-1.2.3/` |
| **DTVM `dtvm`** | DTVM source fork (commits in `VERSIONS.md`) | not vanilla WAMR |

### Build Toolchain (experiment notes / lab notebook)

| Tool | Version / path | Used for |
|------|----------------|----------|
| **rustc** | **≥ 1.85.0** (Wasmtime build) | Wasmtime baseline |
| **LLVM / clang (DTVM build)** | **15.0.0**, e.g. `/opt/clang+llvm-15.0.0-x86_64-linux-gnu-rhel-8.4/...` | Coremark / dtvm builds |
| **LLVM (Wasmer build)** | **18.1.7**, e.g. `...-ubuntu-18.04-...` | Wasmer build |
| **WASI SDK clang** | archived wasm metadata: **11.0.0** (27 cases), **14.0.3** (3 cases) | PolyBench (`benchmarks/polybenchc/BUILD.md`) |
| **WAMR AOT LLVM** | **11.1.0** | `runtimes/wamr-1.2.3/BUILD.md` |
| **Emscripten `em++`** | **not pinned** | `benchmarks/overflow/build.sh`, `-O2` |
| **clang (generic wasm example)** | `--target=wasm32` (no version pin) | md5 example in the experiment notebook |

Coremark self-reports **CLANG 14.0.6** (`-O2`) inside the wasm output — that is the **benchmark program's** compiler, not the host clang used to build DTVM.

### Cache Handling When Comparing Baselines (non-contract)

The archive notes that for Wasmtime / Wasmer comparisons one may `rm -rf ~/.cache/wasmtime`, set `WASMTIME_CACHE=0`, or use `wasmer run --cache-dir=/dev/null`. **No corresponding step was recorded for the contract vm_benchmark.**

---

## Related Documents

| Document | Contents |
|----------|----------|
| [`VERSIONS.md`](VERSIONS.md) | Runtime pins and toolchain summary |
| [`benchmarks/contracts/REPRODUCE.md`](benchmarks/contracts/REPRODUCE.md) | Contract build, run, and measurement methodology |
| [`REPRODUCE.md`](REPRODUCE.md) | Whole-package reproduction skeleton |
77 changes: 77 additions & 0 deletions benchmarks/paper/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
# DTVM Paper Reproduction Package

This directory archives the test cases, benchmark scripts, baseline-runtime build notes, and reference results needed to reproduce every experiment in the Evaluation section of the DTVM paper (`resources/DTVM_paper.pdf`). The material was curated from an experiment archive; the interior layout matches the original package, and every script under `scripts/` treats this directory as its `ROOT`.

Mainline comparison: **DTVM (multipass / lazy)** vs **Wasmtime 45.0.0** vs **Wasmer 7.1.0 (cranelift / singlepass / llvm)**; the EVM-side baseline is **evmone**.

## Layout

```
benchmarks/paper/
├── VERSIONS.md # runtime version pins (paper baseline + current mainline)
├── ENVIRONMENT.md # experiment environment and toolchain inventory
├── REPRODUCE.md # reproduction overview (links to per-topic guides)
├── SOURCES.md # provenance and license notes
├── docs/ # per-topic reproduction guides
│ ├── POLYBENCH_REPRODUCE.md # PolyBench wall-clock
│ ├── POLYBENCH_TTFI_REPRODUCE.md # PolyBench TTFI
│ └── WAPM_TTFI_REPRODUCTION.md # WAPM TTFI (wasmtime 45 / wasmer 7.1 backends)
├── runtimes/ # per-runtime BUILD.md / build.sh / version.txt (no binaries)
├── benchmarks/ # paper workloads
│ ├── contracts/ # contract microbenchmark: sol sources + evm/wasm hex (runner source not included)
│ ├── polybenchc/ # PolyBench/C 4.2.1, all 30 cases as wasm + WASI rebuild script
│ ├── wapm/ # 21 WAPM wasm files (paper's 20 cases + fortune) + confs + SHA256SUMS
│ ├── overflow/ # integer-overflow (swap-style) + fib(30) 5-way: C++ source + wasm + scripts
│ └── fib/ # standalone fib / md5 (C/Solidity sources + wasm)
└── scripts/ # benchmark drivers and table generators (incl. testsuite runner)

`raw_data/` is created by the scripts at run time and is gitignored; this package ships cases and methodology only, no result data.
```

## Paper Results → Directory Map

| Paper result | Experiment | Workloads | Entry point |
|---|---|---|---|
| Fig. "Ethereum Ecosystem" (Fig. 5) | DTVM vs evmone contract latency | `benchmarks/contracts/` (Merkle-MiMC, GenerativeNFT, ERC1155, Counter, ERC20/721, UniswapV2, fib) | `benchmarks/contracts/REPRODUCE.md` |
| Fig. "JIT improvement over Interp" (Fig. 6) | DTVM JIT vs interpreter | `benchmarks/contracts/` + `benchmarks/fib/` | same + `benchmarks/fib/` |
| Fig. "general-purpose Wasm VMs" (Fig. 7) | fib(30) / integer-overflow latency | `benchmarks/overflow/`, `benchmarks/fib/` | `benchmarks/overflow/REPRODUCE_fib_overflow_5way.md` |
| Fig. 8(a) + Appendix PolyBench table | PolyBench 30-case latency | `benchmarks/polybenchc/` | `docs/POLYBENCH_REPRODUCE.md` |
| Fig. 8(b) + Appendix PolyBench TTFI | PolyBench time-to-first-invocation | `benchmarks/polybenchc/` | `docs/POLYBENCH_TTFI_REPRODUCE.md` |
| Fig. 8(b) + Appendix WAPM table | WAPM 20-case TTFI | `benchmarks/wapm/` | `docs/WAPM_TTFI_REPRODUCTION.md`, `benchmarks/wapm/WAPM_REPRODUCE.md` |
| Deterministic-execution table | cross-arch stack-overflow determinism | `tests/wast/dwasm/call_stack_exceed.wast` (already in this repo) | `docs/testing/` |
| Machine-code-size table | compiled-code size comparison | `benchmarks/polybenchc/` + `benchmarks/fib/` | per-topic docs |


## Quick Start

```bash
cd benchmarks/paper
export ROOT=$(pwd)
# 1. Build DTVM at the repo root:
# cmake -B build -DZEN_ENABLE_MULTIPASS_JIT=ON -DLLVM_DIR=<llvm15>/lib/cmake/llvm && cmake --build build
# Scripts default to <repo>/build/dtvm; override via the DTVM env var.
# 2. Install baseline runtimes (see docs/<TOPIC>_REPRODUCE.md and runtimes/<rt>/BUILD.md).
# 3. PolyBench wall-clock
./scripts/setup_polybench_case.sh
./scripts/run_polybench.sh dtvm
./scripts/bench_polybench_timing.sh dtvm multipass
# 4. PolyBench TTFI
WASMTIME_CACHE=0 ./scripts/run_polybench_ttfi_round.sh 1
# 5. WAPM TTFI
WASMTIME_CACHE=0 ./scripts/run_wapm_v45_v71_round.sh 3
# 6. overflow / fib(30)
cd benchmarks/overflow && bash run_dtvm.sh && bash run_wasmtime45.sh
```

## Large Files and Provenance

- `*.wasm`, `*.expected`, and `*.png` are managed by **Git LFS** (see the root `.gitattributes`).
- Per-suite provenance and integrity: `benchmarks/*/MANIFEST.md`, `benchmarks/*/SHA256SUMS`, `SOURCES.md`.
- Runtime binaries are not committed: `runtimes/<rt>/` keeps only `BUILD.md` + `build.sh` + `version.txt` (`out/` and `build/` are gitignored).

## Not Migrated / Archive Gaps

- Full intermediate data (multi-round timestamped runs, per-runtime CSVs, SVG comparisons) and the raw per-iteration contract timing samples (`.time`) are not included.
- The 9 non-paper WAPM cases and the libsodium (72) / wabench (22) suites were not migrated (see `benchmarks/README.md`).
- The contract microbenchmark runner (`test_vmbench`) builds only inside the contract-testbed repository; the sol→wasm / sol→evm compile commands were not recorded in the archive (see `benchmarks/contracts/README.md`, `SOURCES.md`).
- Other gaps are listed in `SOURCES.md` and `benchmarks/wapm/WAPM_REPRODUCE.md` §2.
152 changes: 152 additions & 0 deletions benchmarks/paper/REPRODUCE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,152 @@
# Reproduction Guide

> This package lives at `benchmarks/paper/` in the DTVM repository. All relative paths below (`./scripts/...`, `benchmarks/...`) are rooted at this directory — `cd benchmarks/paper` first.

This document gives the install and run entry points for the three main experiment groups; details live in `docs/<TOPIC>_REPRODUCE.md` and `runtimes/<rt>/BUILD.md`.

## Requirements

- OS: **Linux x86_64**
- Common tools: `cmake` ≥ 3.20, `ninja`, `curl`, `python3`, `taskset`
- DTVM build: GCC C++17, **LLVM 15** (multipass JIT)
- Wasmtime 45 source build: **Rust ≥ stable** (via `rustup`, `source ~/.cargo/env`)
- Wasmer 7.1 LLVM backend: **LLVM 21** (`LLVM_SYS_211_PREFIX` → `LLVM-21.1.x-Linux-X64/`)
- overflow / fib: **Emscripten** (wasm compilation)

Contract-bench machine (from experiment notes): compute-optimized **c7 8C16G**, Xeon **8369B**, **Turbo Boost off**, observed **cpu MHz ≈ 2699.998**. The full 12-item environment/toolchain checklist is in [`ENVIRONMENT.md`](ENVIRONMENT.md).

## Step 1 — Install Runtimes

Each runtime directory keeps only `BUILD.md` + `build.sh` + `version.txt`; binaries are installed to `runtimes/<rt>/out/` (gitignored).

### DTVM (dtvm)

```bash
# Paper PolyBench / WAPM wall-clock: commit 882c83155
# TTFI mainline: main HEAD (includes "Total compilation time" instrumentation)
# overflow / fib: newest fastest commit (e.g. e532db3e2)
```

See [`runtimes/dtvm/BUILD.md`](runtimes/dtvm/BUILD.md).

### Wasmtime

| Purpose | Version | Path |
|---------|---------|------|
| Paper PolyBench wall-clock | **31.0.0** | `runtimes/wasmtime-31.0.0/` (includes `patches/ttfi-report.patch`) |
| TTFI / overflow / fib mainline | **45.0.0** | `runtimes/wasmtime-45.0.0/` (source tree must include the TTFI instrumentation) |

```bash
# 31.0.0: official prebuilt tarball, or build from source
cd runtimes/wasmtime-31.0.0 && ./build.sh all
# 45.0.0: build from an instrumented source tree
WASMER_SRC=/path/to/wasmtime-45.0.0 runtimes/wasmtime-45.0.0/build.sh
```

### Wasmer

| Purpose | Version | Path |
|---------|---------|------|
| Paper PolyBench `wasmer llvm` column | **5.0.4** | `runtimes/wasmer-5.0.4/` |
| TTFI / overflow / fib mainline | **7.1.0** | `runtimes/wasmer-7.1.0/` (cranelift + singlepass + llvm) |

```bash
# 5.0.4: cranelift + singlepass + llvm (llvm backend needs LLVM 18)
cd runtimes/wasmer-5.0.4 && ./build.sh
# 7.1.0: cranelift + singlepass (default); llvm needs LLVM 21
LLVM_SYS_211_PREFIX=/path/to/LLVM-21.1.8-Linux-X64 \
WASMER_SRC=/path/to/wasmer-7.1.0 \
runtimes/wasmer-7.1.0/build.sh
```

### WAMR 1.2.3 (optional, paper interpreter baseline)

```bash
cd runtimes/wamr-1.2.3 && ./build.sh all
```

## Step 2 — Main Experiment Groups

### 2.1 PolyBench wall-clock (30 cases)

```bash
./scripts/setup_polybench_case.sh

# Correctness (once per runtime)
./scripts/run_polybench.sh dtvm
./scripts/run_polybench.sh wasmtime

# Timing CSV (1 warmup + 3 timed runs)
./scripts/bench_polybench_timing.sh dtvm multipass # mainline
./scripts/bench_polybench_timing.sh dtvm lazy
./scripts/bench_polybench_timing.sh wasmtime default
./scripts/bench_polybench_timing.sh wasmer llvm
```

Full procedure, command-line differences, and pitfalls when comparing against `dtvm_polybench.xlsx`: [`docs/POLYBENCH_REPRODUCE.md`](docs/POLYBENCH_REPRODUCE.md).

### 2.2 PolyBench TTFI (30 cases)

DTVM `multipass + lazy` vs **Wasmtime 45.0.0** vs **Wasmer 7.1.0 (cranelift / singlepass / llvm)**; the metric is the `Total compilation time` printed on stdout.

```bash
export DTVM=$PWD/runtimes/dtvm_main/dtvm
export WASMTIME=$PWD/runtimes/wasmtime-45.0.0/out/wasmtime
export WASMER=$PWD/runtimes/wasmer-7.1.0/out/bin/wasmer
export WASMTIME_CACHE=0
./scripts/run_polybench_ttfi_round.sh 1
# → raw_data/benchs/polybench_ttfi_<timestamp>/ + polybench_ttfi_compare_latest.{md,csv}
```

Details: [`docs/POLYBENCH_TTFI_REPRODUCE.md`](docs/POLYBENCH_TTFI_REPRODUCE.md).

### 2.3 WAPM TTFI + wall-clock (20 cases)

```bash
export DTVM=/path/to/DTVM/build/dtvm
export WASMTIME=$PWD/runtimes/wasmtime-45.0.0/out/wasmtime
export WASMER=$PWD/runtimes/wasmer-7.1.0/out/bin/wasmer
export WASMTIME_CACHE=0
./scripts/run_wapm_v45_v71_round.sh 3
# → raw_data/benchs/wapm_v45_v71_<timestamp>/ + wapm_v45_v71_latency_table.{md,csv}
```

Details: [`docs/WAPM_TTFI_REPRODUCTION.md`](docs/WAPM_TTFI_REPRODUCTION.md) and [`benchmarks/wapm/WAPM_REPRODUCE.md`](benchmarks/wapm/WAPM_REPRODUCE.md).

### 2.4 Integer overflow + fib(30) (5-way)

Only the **newest fastest** DTVM commit is used; the other four columns are Wasmtime 45 / Wasmer cranelift / singlepass / llvm.

```bash
cd benchmarks/overflow
# Compile the wasm (emscripten) — archived artifacts already exist, so this is skippable
bash build.sh
# 5-way benchmark
bash run_dtvm.sh # DTVM
bash run_wasmtime45.sh # Wasmtime 45 + all three Wasmer backends
```

Full command lines, the 5-rep median methodology, and per-runtime wasm variants: [`benchmarks/overflow/REPRODUCE_fib_overflow_5way.md`](benchmarks/overflow/REPRODUCE_fib_overflow_5way.md).

## Step 3 — Collect Raw Data

Benchmark output is written under `raw_data/benchs/` (generated at run time; gitignored — no result data is shipped).

## Contract Microbenchmark (ERC20 / Uniswap / Merkle Proof / Generative NFT / …)

Sources and hex artifacts are archived under `benchmarks/contracts/`. The runner builds only inside the contract-testbed repository:

```bash
sh build.sh release.vmbench 4
cd test_run
LD_LIBRARY_PATH=. ./test_vmbench --gtest_filter=VMBenchTest.Erc20Evm
```

Newer cases:

```bash
LD_LIBRARY_PATH=. ./test_vmbench --gtest_filter=VMBenchTest.MerkleProofEvm:VMBenchTest.SolMerkleProofWasm
LD_LIBRARY_PATH=. ./test_vmbench --gtest_filter=VMBenchTest.GenerativeNFTEvm:VMBenchTest.SolGenerativeNFTWasm
```

See `benchmarks/contracts/REPRODUCE.md`.
Loading
Loading