Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
133 changes: 133 additions & 0 deletions docs/changes/2026-07-29-evm-stack-boundary-batch/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
# Change: Batch EVM runtime-stack boundary access

- **Status**: Validated
- **Date**: 2026-07-29
- **Tier**: Full

## Overview

Add a first stage of selective stack dematerialization for the EVM
multipass frontend. Non-lifted block entries load their required live-in values
with one batch address calculation and one depth update. Non-lifted exits store
their bottom-to-top logical stack with one top/size update.

Batching is the standard runtime-stack boundary lowering. Full operand-stack
SSA remains unchanged and takes priority, so batching only handles non-lifted
blocks in SSA builds and all block boundaries when SSA is disabled.

## Contract

`EVMMirBuilder` exposes three frontend-internal operations:

- `peekStackBatch(count, skipTop)` returns operands in bottom-to-top order and
leaves runtime depth unchanged;
- `dropStackBatch(count)` updates runtime top and size once;
- `pushStackBatch(values)` stores bottom-to-top values and updates runtime top
and size once.

Empty batches are strict no-ops. The operations do not change `EVMInstance`,
EVMC ABI, gas accounting, exception behavior, or dynamic-jump dispatch.

## Validation

Build and frontend test commands:

```bash
cmake -S . -B build-batch \
-DCMAKE_BUILD_TYPE=Release \
-DZEN_ENABLE_EVM=ON \
-DZEN_ENABLE_MULTIPASS_JIT=ON \
-DZEN_ENABLE_SINGLEPASS_JIT=OFF \
-DZEN_ENABLE_SPEC_TEST=ON \
-DZEN_ENABLE_EVM_STACK_SSA_LIFT=ON \
-DZEN_ENABLE_EVM_MEMORY_PLAN_FRAMEWORK=ON \
-DLLVM_DIR=/path/to/llvm-15/lib/cmake/llvm
cmake --build build-batch --target evmJitFrontendTests evmDifferentialTests \
evmStateTests
./build-batch/evmJitFrontendTests
./build-batch/evmDifferentialTests
```

Before the experimental gate was removed, validation used two Release builds
from the same source revision:

- A: V111 with boundary batching disabled;
- B: V111 with boundary batching enabled.

The change was also rebuilt as V011 (SSA disabled) in both A/B modes. Results:

| Gate | Result |
| --- | --- |
| V011 frontend | A/B common suite: 123/123; B-only batch tests: 2/2 |
| V111 frontend | A/B common suite: 123/123; B-only batch tests: 2/2 |
| V011 differential | A/B: 73/73 |
| V111 differential | A/B: 73/73 |
| state tests | B: 1798/1798 |
| forced JIT-to-interpreter fallback | B: 8/8 |
| Osaka transaction-exact replay | B: 100/100, 29 code hashes |

The exact replay used the fixture-aware state-test host, which restores complete
prestate before each transaction. Its result is
`/tmp/dtvm-pr1-b-correctness-100.json` on the measurement host, SHA-256
`8c2e8bbcd6b68010f49800f963ee54dfef23ae6685e8604a73f1abbc40c50c3d`.

The synthetic 16-slot, 8-boundary stress case reduced MIR instructions from
5147 to 3131 (-39.2%) and CgIR instructions from 4128 to 2614 (-36.7%).
Production compiler observation over the 29 real code hashes recorded 7103
batch loads/drops for 22004 slots and 7383 batch stores for 26359 slots. Top
and size were each updated 14486 times, exactly once per drop or store batch.
These coverage counters were collected with the optional compiler-observation
instrumentation from PR #587; that instrumentation is intentionally not part
of this change and is not required by the optimization.

Formal Osaka performance used CPU 24, 29 unique code hashes, 12 fresh-process
cold rounds, three warmups per variant, and 100 ms hot calibration. S0 denotes
A and S1 denotes B in the result file:

| Metric | B relative to A | 95% bootstrap CI |
| --- | ---: | ---: |
| JIT compilation, geometric mean | -2.645% | [-3.202%, -2.098%] |
| Emitted code, geometric mean | -0.168% | [-0.216%, -0.122%] |
| Hot execution, geometric mean | -0.142% | [-0.375%, +0.082%] |
| Frequency-weighted hot execution | -0.243% | [-0.626%, +0.087%] |

The result is `/tmp/dtvm-pr1-ab-formal-29x12.json`, SHA-256
`c9c71d8c9ccf7b594ce9bcf28d64dacf0cc4b52f2a080055dccf80c87c6f5f6d`.
The compiler-observation result is
`/tmp/dtvm-pr1-ab-observation-29.json`, SHA-256
`f38d5d865554ea829de3fb3a00df4d705a906d786fe859bbb0e826bed0dd834e`.
Both compile/code-size and hot-execution regression gates pass.

Reproduce exact correctness and paired performance with the existing
fixture-aware replay tools:

```bash
python3 tools/run_replay_correctness.py \
--executable build-b-v111/evmStateTests \
--fixture-manifest /path/to/fixtures-100-v2-manifest.json \
--fixture-dir /path/to/fixtures-100-v2 \
--output /tmp/dtvm-pr1-b-correctness-100.json \
--raw-dir /tmp/dtvm-pr1-b-correctness-100-raw \
--mode multipass --revision Osaka --cpu 24

python3 tools/run_ssa_replay_exact_performance.py \
--s0 build-a-v111/evmStateTests \
--s1 build-b-v111/evmStateTests \
--performance-set /path/to/performance-set/manifest.json \
--fixture-dir /path/to/29-fixtures \
--output /tmp/dtvm-pr1-ab-formal-29x12.json \
--raw-output-dir /tmp/dtvm-pr1-ab-formal-29x12-raw \
--revision Osaka --cpu 24 --rounds 12 \
--warmup-iterations 3 --target-hot-ms 100 \
--bootstrap-samples 10000 --seed 20260729
```

## Risks

- A wrong base offset can reverse U256 or stack-slot order. Batch tests cover
empty, 1-, 16-, and 17-slot shapes; differential and exact replay remain
mandatory.
- Batching reduces address and top/size maintenance but not the four limb
loads/stores per U256 value.
- Observation counters are compile-time structural evidence and are not a
substitute for hot-execution measurement.
1 change: 1 addition & 0 deletions docs/changes/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,7 @@ Typical triggers:
| 2026-05-13 | [evm-ngram-macro-ops](2026-05-13-evm-ngram-macro-ops/README.md) | Implemented | Full | Initial EVM n-gram macro-op lowering and specialized keccak helpers for multipass JIT |
| 2026-07-21 | [evm-memory-alias-and-expansion](2026-07-21-evm-memory-alias-and-expansion/README.md) | Implemented | Full | Stronger memory alias proofs, wider precheck/expansion coverage, DSE, load forwarding, grouping, and MCOPY roadmap |
| 2026-07-28 | [ssa-shared-dynamic-dispatch](2026-07-28-ssa-shared-dynamic-dispatch/README.md) | Implemented | Light | Share unfiltered full-table dynamic dispatch in stack-SSA builds |
| 2026-07-29 | [evm-stack-boundary-batch](2026-07-29-evm-stack-boundary-batch/README.md) | Validated | Full | Batch runtime-stack loads, drops, and stores at non-lifted EVM block boundaries |
| 2026-08-22 | [evm-frontier-osaka-conformance](2026-08-22-evm-frontier-osaka-conformance/README.md) | Accepted | Full | Pinned dual-engine EVM conformance from Frontier through Osaka |

Each active proposal lives in its own subdirectory. Browse `docs/changes/*/README.md`
Expand Down
12 changes: 12 additions & 0 deletions docs/modules/compiler/spec.md
Original file line number Diff line number Diff line change
Expand Up @@ -212,6 +212,18 @@ and destination ends for expansion elision while retaining the original
- Provide bytecode, gas chunk end/cost arrays for chunk-based metering
- Use register to hold gas when `ZEN_ENABLE_EVM_GAS_REGISTER` is enabled

### EVM Stack Boundary Batching

Non-lifted EVM block boundaries use batched runtime-stack access. Entry loads
use `peekStackBatch` followed by one `dropStackBatch`; exit materialization uses
one `pushStackBatch`. Batch vectors are always bottom-to-top. Empty batches emit
no MIR and do not update runtime stack state. Full stack-SSA blocks retain the
existing entry-state and phi protocol.

The runtime stack remains the complete authoritative representation in this
stage. Every non-lifted exit still writes all logical values, and dynamic
dispatch, fallback, gas, and runtime ABI behavior are unchanged.

### EVM Stack SSA Lift Safety

- `ZEN_ENABLE_EVM_STACK_SSA_LIFT` permits compatible EVM operand-stack values
Expand Down
26 changes: 10 additions & 16 deletions src/action/evm_bytecode_visitor.h
Original file line number Diff line number Diff line change
Expand Up @@ -1039,9 +1039,7 @@ template <typename IRBuilder> class EVMByteCodeVisitor {
// missed underflow traps in later blocks.
spillTrackedStackPreservingPrefix(Values, /*PrefixDepth=*/0);
} else {
for (const Operand &Opnd : Values) {
Builder.stackPush(Opnd);
}
Builder.pushStackBatch(Values);
}
}
InDeadCode = true;
Expand Down Expand Up @@ -1292,27 +1290,23 @@ template <typename IRBuilder> class EVMByteCodeVisitor {

CurrentBlockLifted = false;
int32_t TotalPopSize = -BlockInfo.MinPopHeight;
EvalStack ReverseStack;
// Refine each popped Operand's ValueRange from analyzer-computed entry
// ranges so u64-narrow fast paths fire on values flowing through CFG
// joins (see EVMRangeAnalyzer /
// docs/changes/2026-05-07-value-range-cfg-join). EntryStackRanges[0] is the
// bottom of entry stack; pop order is top-first.
const auto &EntryRanges = BlockInfo.EntryStackRanges;
const int32_t EntryTopIdx = static_cast<int32_t>(EntryRanges.size()) - 1;
int32_t PopIter = 0;
while (TotalPopSize > 0) {
Operand Opnd = Builder.stackPop();
const int32_t SlotIdx = EntryTopIdx - PopIter;
std::vector<Operand> EntryValues =
Builder.peekStackBatch(static_cast<uint32_t>(TotalPopSize));
Builder.dropStackBatch(static_cast<uint32_t>(TotalPopSize));
const int32_t FirstEntrySlot =
static_cast<int32_t>(EntryRanges.size()) - TotalPopSize;
for (int32_t Index = 0; Index < TotalPopSize; ++Index) {
Operand &Opnd = EntryValues[static_cast<size_t>(Index)];
const int32_t SlotIdx = FirstEntrySlot + Index;
if (SlotIdx >= 0 && SlotIdx < static_cast<int32_t>(EntryRanges.size())) {
Opnd.setRange(EntryRanges[SlotIdx]);
Opnd.setRange(EntryRanges[static_cast<size_t>(SlotIdx)]);
}
ReverseStack.push(Opnd);
++PopIter;
--TotalPopSize;
}
while (!ReverseStack.empty()) {
Operand Opnd = ReverseStack.pop();
Stack.push(Opnd);
}
}
Expand Down
96 changes: 96 additions & 0 deletions src/compiler/evm_frontend/evm_mir_compiler.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1110,6 +1110,102 @@ typename EVMMirBuilder::Operand EVMMirBuilder::stackPop() {
return Operand(PopComponents, EVMType::UINT256);
}

std::vector<EVMMirBuilder::Operand>
EVMMirBuilder::peekStackBatch(uint32_t Count, uint32_t SkipTop) {
std::vector<Operand> Values;
if (Count == 0) {
return Values;
}

const uint64_t AddressedSlots =
static_cast<uint64_t>(Count) + static_cast<uint64_t>(SkipTop);
ZEN_ASSERT(AddressedSlots <= 1024 &&
"runtime stack batch peek exceeds the EVM stack");

MType *I64Type = EVMFrontendContext::getMIRTypeFromEVMType(EVMType::UINT64);
MPointerType *U64PtrType = MPointerType::create(Ctx, Ctx.I64Type);
MInstruction *StackTopInt = getInstanceStackTopInt();
MInstruction *BaseOffset =
createIntConstInstruction(I64Type, AddressedSlots * 32ULL);
MInstruction *BatchBase = createInstruction<BinaryInstruction>(
false, OP_sub, I64Type, StackTopInt, BaseOffset);
MInstruction *BatchPtr = createInstruction<ConversionInstruction>(
false, OP_inttoptr, U64PtrType, BatchBase);

Values.reserve(Count);
for (uint32_t Slot = 0; Slot < Count; ++Slot) {
U256Inst Components = {};
for (size_t Limb = 0; Limb < EVM_ELEMENTS_COUNT; ++Limb) {
const int32_t Offset =
static_cast<int32_t>(static_cast<uint64_t>(Slot) * 32ULL +
static_cast<uint64_t>(Limb) * 8ULL);
MInstruction *LoadInstr = createInstruction<LoadInstruction>(
false, I64Type, BatchPtr, 1, nullptr, Offset);
Variable *ValVar = storeInstructionInTemp(LoadInstr, I64Type);
Components[Limb] = loadVariable(ValVar);
}
Values.emplace_back(Components, EVMType::UINT256);
}

return Values;
}

void EVMMirBuilder::dropStackBatch(uint32_t Count) {
if (Count == 0) {
return;
}
ZEN_ASSERT(Count <= 1024 && "runtime stack batch drop exceeds the EVM stack");

MType *I64Type = EVMFrontendContext::getMIRTypeFromEVMType(EVMType::UINT64);
MInstruction *StackBytes =
createIntConstInstruction(I64Type, static_cast<uint64_t>(Count) * 32ULL);
MInstruction *StackTopInt = getInstanceStackTopInt();
MInstruction *StackSize = loadVariable(StackSizeVar);
MInstruction *NewTop = createInstruction<BinaryInstruction>(
false, OP_sub, I64Type, StackTopInt, StackBytes);
MInstruction *NewSize = createInstruction<BinaryInstruction>(
false, OP_sub, I64Type, StackSize, StackBytes);
createInstruction<DassignInstruction>(true, &(Ctx.VoidType), NewTop,
StackTopVar->getVarIdx());
createInstruction<DassignInstruction>(true, &(Ctx.VoidType), NewSize,
StackSizeVar->getVarIdx());
}

void EVMMirBuilder::pushStackBatch(const std::vector<Operand> &Values) {
if (Values.empty()) {
return;
}
ZEN_ASSERT(Values.size() <= 1024 &&
"runtime stack batch push exceeds the EVM stack");

MType *I64Type = EVMFrontendContext::getMIRTypeFromEVMType(EVMType::UINT64);
MPointerType *U64PtrType = MPointerType::create(Ctx, Ctx.I64Type);
MInstruction *StackTopInt = getInstanceStackTopInt();
MInstruction *StackTopPtr = createInstruction<ConversionInstruction>(
false, OP_inttoptr, U64PtrType, StackTopInt);

for (size_t Slot = 0; Slot < Values.size(); ++Slot) {
U256Inst Components = extractU256Operand(Values[Slot]);
for (size_t Limb = 0; Limb < EVM_ELEMENTS_COUNT; ++Limb) {
const int32_t Offset = static_cast<int32_t>(Slot * 32ULL + Limb * 8ULL);
createInstruction<StoreInstruction>(true, &Ctx.VoidType, Components[Limb],
StackTopPtr, Offset);
}
}

MInstruction *StackSize = loadVariable(StackSizeVar);
MInstruction *StackBytes = createIntConstInstruction(
I64Type, static_cast<uint64_t>(Values.size()) * 32ULL);
MInstruction *NewTop = createInstruction<BinaryInstruction>(
false, OP_add, I64Type, StackTopInt, StackBytes);
MInstruction *NewSize = createInstruction<BinaryInstruction>(
false, OP_add, I64Type, StackSize, StackBytes);
createInstruction<DassignInstruction>(true, &(Ctx.VoidType), NewTop,
StackTopVar->getVarIdx());
createInstruction<DassignInstruction>(true, &(Ctx.VoidType), NewSize,
StackSizeVar->getVarIdx());
}

void EVMMirBuilder::stackSet(int32_t IndexFromTop, Operand SetValue) {
// This set element to stack with index from top
U256Inst SetComponents = extractU256Operand(SetValue);
Expand Down
3 changes: 3 additions & 0 deletions src/compiler/evm_frontend/evm_mir_compiler.h
Original file line number Diff line number Diff line change
Expand Up @@ -353,6 +353,9 @@ class EVMMirBuilder final {

void stackPush(Operand PushValue);
Operand stackPop();
std::vector<Operand> peekStackBatch(uint32_t Count, uint32_t SkipTop = 0);
void dropStackBatch(uint32_t Count);
void pushStackBatch(const std::vector<Operand> &Values);

void stackSet(int32_t IndexFromTop, Operand SetValue);
Operand stackGet(int32_t IndexFromTop);
Expand Down
61 changes: 61 additions & 0 deletions src/tests/evm_differential_tests.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -841,6 +841,67 @@ TEST(EVMRangeDifferential, DeadUnderResolvedFallthroughMatchesInterpreter) {
Bytecode, {}));
}

TEST(EVMStackBoundaryDifferential,
SixteenSlotForwardBoundaryStressMatchesInterpreter) {
std::vector<uint8_t> Bytecode;
auto AppendPush2 = [&](uint16_t Value) {
Bytecode.push_back(0x61);
Bytecode.push_back(static_cast<uint8_t>(Value >> 8));
Bytecode.push_back(static_cast<uint8_t>(Value));
};

constexpr uint32_t Slots = 16;
constexpr uint32_t Boundaries = 8;
for (uint32_t Slot = 0; Slot < Slots; ++Slot) {
AppendPush2(static_cast<uint16_t>(Slot + 1));
}
for (uint32_t Boundary = 0; Boundary < Boundaries; ++Boundary) {
const uint16_t TargetPC = static_cast<uint16_t>(Bytecode.size() + 4);
AppendPush2(TargetPC);
Bytecode.push_back(0x56); // JUMP
Bytecode.push_back(0x5b); // JUMPDEST
Bytecode.insert(Bytecode.end(), Slots, 0x50); // POP x16
for (uint32_t Slot = 0; Slot < Slots; ++Slot) {
AppendPush2(static_cast<uint16_t>((Boundary + 1) * Slots + Slot + 1));
}
}
Bytecode.push_back(0x00); // STOP

EXPECT_TRUE(expectInterpMatchesMultipass(
"sixteen_slot_forward_boundary_stress", Bytecode, {}));
}

TEST(EVMStackBoundaryDifferential,
PopDup16AndSwap16AcrossBoundaryMatchInterpreter) {
std::vector<uint8_t> Bytecode;
for (uint8_t Value = 1; Value <= 17; ++Value) {
Bytecode.push_back(0x60); // PUSH1
Bytecode.push_back(Value);
}
Bytecode.push_back(0x60); // PUSH1 target
Bytecode.push_back(static_cast<uint8_t>(Bytecode.size() + 2));
Bytecode.push_back(0x56); // JUMP
Bytecode.push_back(0x5b); // JUMPDEST
Bytecode.push_back(0x8f); // DUP16
Bytecode.push_back(0x9f); // SWAP16
Bytecode.push_back(0x50); // POP
Bytecode.insert(Bytecode.end(),
{0x5f, 0x52, 0x60, 0x20, 0x5f, 0xf3}); // return top word

EXPECT_TRUE(expectInterpMatchesMultipass("pop_dup16_swap16_across_boundary",
Bytecode, {}));
}

TEST(EVMStackBoundaryDifferential, UnderflowAndOverflowMatchInterpreter) {
EXPECT_TRUE(expectInterpStatusMatchesMultipass(
"batch_boundary_underflow", {0x50, 0x00}, EVMC_STACK_UNDERFLOW));

std::vector<uint8_t> Overflow(1025, 0x5f); // PUSH0 x1025
Overflow.push_back(0x00);
EXPECT_TRUE(expectInterpStatusMatchesMultipass(
"batch_boundary_overflow", Overflow, EVMC_STACK_OVERFLOW));
}

TEST(EVMLiftedStackMerge,
JumpiFallthroughSharedJumpDestMatchesInterpreterOnBothEdges) {
const std::vector<uint8_t> Bytecode = {
Expand Down
Loading
Loading