You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On x64, when a method's prolog must zero exactly four int-sized stack slots that happen to form one dense 16-byte range, the JIT emits xor reg, reg plus two pointer-sized stores (12 bytes / 3 instructions) instead of the xorps xmm + single 16-byte store sequence it already knows how to emit (10 bytes / 2 instructions).
The reason is a phase-ordering limitation rather than a deliberate tuning choice. CodeGen::genCheckUseBlockInit runs before frame offsets are assigned, so it can only count int-sized slots; its AMD64 threshold is genInitStkLclCnt > 4. That excludes exactly the 16-byte case, while a 24-byte local (6 slots) does get block init. By the time CodeGen::genZeroInitFrame runs the final offsets are known, so the dense-16-byte case can be recognized there and routed to the existing genZeroInitFrameUsingBlockInit path.
This is a common shape — any method with a single 16-byte .locals init struct, or two GC-pointer locals that land adjacently.
Method size 33 → 31 bytes. No frame-layout change, no frame growth, no new callee-save: xmm4 (Windows) / xmm8 (SysV) is the same volatile scratch the existing block-init path already uses. Unaligned vmovdqu and aligned vmovdqa are the same length, so an unaligned 16-byte range still wins.
The saving grows on large frames, where each scalar store pays disp32+SIB. Example from the corpus, Microsoft.CodeAnalysis.CSharp.Binder:BindSimpleProgramCompilationUnit(...) (FullOpts, 1312 → 1306 bytes):
SuperPMI asmdiffs, windows-x64, Release JIT pair, all 12 default collections (JIT-EE GUID fbbaf45f-5b0e-4767-b962-5084c8caae77), base commit 4856f0c16f89d8cc625a62ee7bab7ed4afe14f9d:
Collection
Delta bytes
%
Ctx w/ diffs
Impr
Regr
aspire.nativeaot
−4,207
−0.03
1,127
1,127
0
aspnet2.run
−2,142
−0.01
1,624
1,622
2
benchmarks.run
−1,139
−0.01
654
652
2
benchmarks.run_pgo
−3,381
−0.01
3,166
3,165
0
benchmarks.run_pgo_optrepeat
−1,135
−0.01
653
651
2
coreclr_tests.run
−30,557
−0.01
27,513
27,512
1
libraries.crossgen2
−8,901
−0.02
4,908
4,907
1
libraries.pmi
−8,121
−0.01
4,727
4,721
6
libraries_tests.run
−50,901
−0.01
47,579
47,575
3
libraries_tests_no_tiered_compilation.run
−21,808
−0.01
11,953
11,949
4
realworld.run
−1,349
−0.01
846
846
0
smoke_tests.nativeaot
−1,979
−0.03
625
625
0
Total
−135,620
−0.012%
105,375
105,352
21
Zero failing compiles in every collection, base and diff. No new asserts. Missing-context counts are identical in base and diff everywhere except libraries_tests.run, where the diff JIT has fewer (29 vs 38).
Known regressions. All 21 regressions are exactly +1 byte, and all have the same cause: when the integer zero register is already live for other reasons the xor is not actually saved, and on an rbp frame with disp8 the two 4-byte scalar stores (8 bytes) beat vxorps + vmovdqu (9 bytes). Representative case, Microsoft.Extensions.Logging.LoggerFactory:RefreshFilters(...) (Tier1, 279 → 280):
base: xoredi,edi ; edi is also the zero source for later usesmov qword ptr [rbp-0x48],rdi ; 4mov qword ptr [rbp-0x40],rdi ; 4diff: vxorps xmm4,xmm4,xmm4 ; 4 (xor edi,edi survives — still needed) vmovdqu xmmword ptr [rbp-0x48],xmm4 ; 5
Suppressing these would need initReg liveness information that is not available at this point in the prolog. 21 bytes against −135,641 bytes of improvements.
Throughput.superpmi.py tpdiff (PIN) on 4 of 12 collections before the session budget was exhausted: aspire.nativeaot +0.00%, aspnet2.run +0.01%, benchmarks.run +0.00%, benchmarks.run_pgo +0.01% of total instructions executed. Both JITs carry superpmi.py's "compiled with native PGO" warning, so sub-0.05% readings are within noise.
Measurement limitations:
PerfScore is unmeasured. Both JITs in this experiment were Release builds and RyuJIT only computes PerfScore under DEBUG/LATE_DISASM, so every collection reports Total PerfScore 0.0 / geomean 1.0. No execution-speed claim is made here; this is a code-size and instruction-count result.
tpdiff covers 4 of 12 collections.
Only windows-x64 collections were replayed. Linux/SysV AMD64 (different scratch register, different calleeRegArgMaskLiveIn, different frame offsets) was not diffed and is the one gap worth closing.
No runtime test suite was run against the prototype (JIT-only build). Correctness evidence is the repro's identical managed output (OK acc=4, exit 100 on both JITs) plus 105,375 replayed contexts with zero failing compiles.
Notes
Scope. AMD64 only in the prototype (#ifdef TARGET_AMD64). x86, arm, arm64, loongarch64, riscv64 and wasm are untouched; their thresholds and block-init helpers are unmodified. arm64 may want an analogous treatment via stp, but that is not evaluated here.
Threshold left alone. Lowering genCheckUseBlockInit's AMD64 genInitStkLclCnt > 4 to > 3 would be the obvious fix but is wrong: it would also admit sparse 4-slot sets whose [untrLclLo, untrLclHi) span covers unrelated locals, zeroing memory that the scalar path does not touch. The density check has to happen after offsets are assigned.
Same bytes zeroed. Requiring genInitStkLclCnt == 4anduntrLclHi - untrLclLo == 16 forces the range to be exactly 16 bytes fully covered by lvMustInit slots. A further guard requires that the scalar path would have written all 16 of those bytes — this excludes the TYP_STRUCT + compInitMem == false case where the scalar path writes only the GC slots (block init would then emit 10 bytes to replace a single 5-byte store). Nothing outside the must-init locals is touched, so no GC-reportable slot changes meaning.
initRegZeroed is already handled.genZeroInitFrameUsingBlockInit leaves *pInitRegZeroed untouched when it does not need the integer zero, so later prolog consumers (GS cookie, generic context arg) correctly re-materialize it. That is pre-existing behavior on the genUseBlockInit path; this only widens which methods take it.
Compiler::fgVarNeedsExplicitZeroInit is deliberately not changed. It consumes genCheckUseBlockInit's threshold to decide which locals can drop an explicit body-level zero init. Since this does not alter which locals are prolog-zeroed (lvMustInit is untouched) — only the instruction sequence used — that coupling stays sound. Revisiting it is a possible follow-up, not part of this.
AVX state. No new kind of instruction is introduced; the implicit zeroing of YMM4/ZMM4 and the AVX-state bookkeeping are identical to the existing block-init path, which already emits this sequence for larger ranges.
diff --git a/src/coreclr/jit/codegen.h b/src/coreclr/jit/codegen.h
index 2020048e2ac..616c1356f87 100644
--- a/src/coreclr/jit/codegen.h+++ b/src/coreclr/jit/codegen.h@@ -577,6 +577,9 @@ protected:
void genZeroInitFrame(int untrLclHi, int untrLclLo, regNumber initReg, bool* pInitRegZeroed);
void genZeroInitFrameUsingBlockInit(int untrLclHi, int untrLclLo, regNumber initReg, bool* pInitRegZeroed);
+#ifdef TARGET_AMD64+ bool genCanZeroInitFrameUsingOneSimdStore(int untrLclHi, int untrLclLo);+#endif
void genReportGenericContextArg(regNumber initReg, bool* pInitRegZeroed);
diff --git a/src/coreclr/jit/codegencommon.cpp b/src/coreclr/jit/codegencommon.cpp
index 5b257d36dc9..c3af59312eb 100644
--- a/src/coreclr/jit/codegencommon.cpp+++ b/src/coreclr/jit/codegencommon.cpp@@ -4197,6 +4197,82 @@ regNumber CodeGen::genGetZeroReg(regNumber initReg, bool* pInitRegZeroed)
#endif // !TARGET_ARM64
}
+#ifdef TARGET_AMD64+//-----------------------------------------------------------------------------+// genCanZeroInitFrameUsingOneSimdStore: Determine whether the prolog zeroing work is a single+// dense 16-byte range that is cheaper to clear with one SIMD store.+//+// Arguments:+// untrLclHi - (Untracked locals High-Offset) The upper bound offset (not inclusive) of the+// range that will be zeroed.+// untrLclLo - (Untracked locals Low-Offset) The lower bound offset of that range.+//+// Return Value:+// True if genZeroInitFrameUsingBlockInit should be used even though genCheckUseBlockInit did+// not set genUseBlockInit.+//+// Notes:+// genCheckUseBlockInit runs before frame offsets are assigned, so it can only look at the+// number of int-sized slots to zero and has to stay conservative. When those slots turn out+// to form one contiguous 16-byte range that the scalar path would store in full, block init+// emits "xorps xmm, xmm" plus a single 16-byte store, which is one instruction and two bytes+// shorter than zeroing an integer register plus two pointer-sized stores.+//+// Ranges the scalar path would only write partially (a struct with non-GC fields when+// compInitMem is false) are excluded, because block init would then be larger than the stores+// it replaces.+//+bool CodeGen::genCanZeroInitFrameUsingOneSimdStore(int untrLclHi, int untrLclLo)+{+ if (genUseBlockInit || (genInitStkLclCnt != (XMM_REGSIZE_BYTES / sizeof(int))) ||+ ((untrLclHi - untrLclLo) != XMM_REGSIZE_BYTES))+ {+ return false;+ }++ // Tally the bytes the scalar path in genZeroInitFrame would store; this mirrors the loop there.+ unsigned scalarZeroedBytes = 0;++ for (unsigned varNum = 0; varNum < m_compiler->lvaCount; varNum++)+ {+ LclVarDsc* varDsc = m_compiler->lvaGetDesc(varNum);++ if (!varDsc->lvMustInit || (varDsc->lvIsInReg() && !varDsc->IsLiveInOutOfHandler()) ||+ m_compiler->lvaIsUnknownSizeLocal(varNum))+ {+ continue;+ }++ if (varDsc->TypeIs(TYP_STRUCT) && !m_compiler->info.compInitMem &&+ (varDsc->lvExactSize() >= TARGET_POINTER_SIZE))+ {+ const unsigned slots = (unsigned)m_compiler->lvaLclStackHomeSize(varNum) / REGSIZE_BYTES;+ ClassLayout* layout = varDsc->GetLayout();++ for (unsigned i = 0; i < slots; i++)+ {+ scalarZeroedBytes += layout->IsGCPtr(i) ? REGSIZE_BYTES : 0;+ }+ }+ else+ {+ scalarZeroedBytes += roundUp(m_compiler->lvaLclStackHomeSize(varNum), (unsigned)sizeof(int));+ }+ }++ assert(regSet.tmpAllFree());+ for (TempDsc* tempThis = regSet.tmpListBeg(); tempThis != nullptr; tempThis = regSet.tmpListNxt(tempThis))+ {+ if (varTypeIsGC(tempThis->tdTempType()))+ {+ scalarZeroedBytes += REGSIZE_BYTES;+ }+ }++ return scalarZeroedBytes == XMM_REGSIZE_BYTES;+}+#endif // TARGET_AMD64+
//-----------------------------------------------------------------------------
// genZeroInitFrame: Zero any untracked pointer locals and/or initialize memory for locspace
//
@@ -4212,6 +4288,13 @@ void CodeGen::genZeroInitFrame(int untrLclHi, int untrLclLo, regNumber initReg,
{
assert(GetEmitter()->emitGeneratingPrologOrFuncletProlog());
+#ifdef TARGET_AMD64+ if (genCanZeroInitFrameUsingOneSimdStore(untrLclHi, untrLclLo))+ {+ genUseBlockInit = true;+ }+#endif+
if (genUseBlockInit)
{
genZeroInitFrameUsingBlockInit(untrLclHi, untrLclLo, initReg, pInitRegZeroed);
On x64, when a method's prolog must zero exactly four int-sized stack slots that happen to form one dense 16-byte range, the JIT emits
xor reg, regplus two pointer-sized stores (12 bytes / 3 instructions) instead of thexorps xmm+ single 16-byte store sequence it already knows how to emit (10 bytes / 2 instructions).The reason is a phase-ordering limitation rather than a deliberate tuning choice.
CodeGen::genCheckUseBlockInitruns before frame offsets are assigned, so it can only count int-sized slots; its AMD64 threshold isgenInitStkLclCnt > 4. That excludes exactly the 16-byte case, while a 24-byte local (6 slots) does get block init. By the timeCodeGen::genZeroInitFrameruns the final offsets are known, so the dense-16-byte case can be recognized there and routed to the existinggenZeroInitFrameUsingBlockInitpath.This is a common shape — any method with a single 16-byte
.locals initstruct, or two GC-pointer locals that land adjacently.Minimal repro
Run with
DOTNET_TieredCompilation=0,DOTNET_ReadyToRun=0,DOTNET_JitDisasm="ProbePair ProbeTriple ProbeQuad ProbeMixed"on windows-x64.Current codegen
Repro:ProbePair()— prolog, 33 bytes total method size:Repro:ProbeTriple()(24 bytes of slots) already takes the block-init path, so the 16-byte case is the odd one out.Expected codegen
Method size 33 → 31 bytes. No frame-layout change, no frame growth, no new callee-save:
xmm4(Windows) /xmm8(SysV) is the same volatile scratch the existing block-init path already uses. Unalignedvmovdquand alignedvmovdqaare the same length, so an unaligned 16-byte range still wins.The saving grows on large frames, where each scalar store pays
disp32+SIB. Example from the corpus,Microsoft.CodeAnalysis.CSharp.Binder:BindSimpleProgramCompilationUnit(...)(FullOpts, 1312 → 1306 bytes):Impact
Target method: 33 → 31 bytes (−2), 3 → 2 prolog instructions, 2 → 1 store uop.
SuperPMI
asmdiffs, windows-x64, Release JIT pair, all 12 default collections (JIT-EE GUIDfbbaf45f-5b0e-4767-b962-5084c8caae77), base commit4856f0c16f89d8cc625a62ee7bab7ed4afe14f9d:Zero failing compiles in every collection, base and diff. No new asserts. Missing-context counts are identical in base and diff everywhere except
libraries_tests.run, where the diff JIT has fewer (29 vs 38).Known regressions. All 21 regressions are exactly
+1byte, and all have the same cause: when the integer zero register is already live for other reasons thexoris not actually saved, and on an rbp frame withdisp8the two 4-byte scalar stores (8 bytes) beatvxorps+vmovdqu(9 bytes). Representative case,Microsoft.Extensions.Logging.LoggerFactory:RefreshFilters(...)(Tier1, 279 → 280):Suppressing these would need
initRegliveness information that is not available at this point in the prolog. 21 bytes against −135,641 bytes of improvements.Throughput.
superpmi.py tpdiff(PIN) on 4 of 12 collections before the session budget was exhausted:aspire.nativeaot +0.00%,aspnet2.run +0.01%,benchmarks.run +0.00%,benchmarks.run_pgo +0.01%of total instructions executed. Both JITs carry superpmi.py's "compiled with native PGO" warning, so sub-0.05% readings are within noise.Measurement limitations:
DEBUG/LATE_DISASM, so every collection reportsTotal PerfScore 0.0/ geomean1.0. No execution-speed claim is made here; this is a code-size and instruction-count result.calleeRegArgMaskLiveIn, different frame offsets) was not diffed and is the one gap worth closing.OK acc=4, exit 100 on both JITs) plus 105,375 replayed contexts with zero failing compiles.Notes
#ifdef TARGET_AMD64). x86, arm, arm64, loongarch64, riscv64 and wasm are untouched; their thresholds and block-init helpers are unmodified. arm64 may want an analogous treatment viastp, but that is not evaluated here.genCheckUseBlockInit's AMD64genInitStkLclCnt > 4to> 3would be the obvious fix but is wrong: it would also admit sparse 4-slot sets whose[untrLclLo, untrLclHi)span covers unrelated locals, zeroing memory that the scalar path does not touch. The density check has to happen after offsets are assigned.genInitStkLclCnt == 4anduntrLclHi - untrLclLo == 16forces the range to be exactly 16 bytes fully covered bylvMustInitslots. A further guard requires that the scalar path would have written all 16 of those bytes — this excludes theTYP_STRUCT+compInitMem == falsecase where the scalar path writes only the GC slots (block init would then emit 10 bytes to replace a single 5-byte store). Nothing outside the must-init locals is touched, so no GC-reportable slot changes meaning.initRegZeroedis already handled.genZeroInitFrameUsingBlockInitleaves*pInitRegZeroeduntouched when it does not need the integer zero, so later prolog consumers (GS cookie, generic context arg) correctly re-materialize it. That is pre-existing behavior on thegenUseBlockInitpath; this only widens which methods take it.Compiler::fgVarNeedsExplicitZeroInitis deliberately not changed. It consumesgenCheckUseBlockInit's threshold to decide which locals can drop an explicit body-level zero init. Since this does not alter which locals are prolog-zeroed (lvMustInitis untouched) — only the instruction sequence used — that coupling stays sound. Revisiting it is a possible follow-up, not part of this.rep stosd; Use simd for small prolog zeroing (ia32/x64) #32442 was closed in its favor. This is the opposite end of the same threshold and does not overlap with either.Prototype patch
Experimental patch
Note
This issue was generated with GitHub Copilot.