Skip to content

JIT: x64 prolog zeroing of a dense 16-byte range should use one SIMD store instead of two scalar stores #134907

Description

@AndyAyersMS

On x64, when a method's prolog must zero exactly four int-sized stack slots that happen to form one dense 16-byte range, the JIT emits xor reg, reg plus two pointer-sized stores (12 bytes / 3 instructions) instead of the xorps xmm + single 16-byte store sequence it already knows how to emit (10 bytes / 2 instructions).

The reason is a phase-ordering limitation rather than a deliberate tuning choice. CodeGen::genCheckUseBlockInit runs before frame offsets are assigned, so it can only count int-sized slots; its AMD64 threshold is genInitStkLclCnt > 4. That excludes exactly the 16-byte case, while a 24-byte local (6 slots) does get block init. By the time CodeGen::genZeroInitFrame runs the final offsets are known, so the dense-16-byte case can be recognized there and routed to the existing genZeroInitFrameUsingBlockInit path.

This is a common shape — any method with a single 16-byte .locals init struct, or two GC-pointer locals that land adjacently.

Minimal repro

using System;
using System.Runtime.CompilerServices;

public static class Repro
{
    public struct Pair   { public object First; public object Second; }
    public struct Triple { public object First; public object Second; public object Third; }
    public struct Quad   { public object First; public object Second; public object Third; public object Fourth; }
    public struct Mixed  { public object First; public int Second; }

    [MethodImpl(MethodImplOptions.NoInlining)] public static object ReadPair(ref Pair p)     => p.First ?? p.Second;
    [MethodImpl(MethodImplOptions.NoInlining)] public static object ReadTriple(ref Triple t) => t.First ?? t.Second ?? t.Third;
    [MethodImpl(MethodImplOptions.NoInlining)] public static object ReadQuad(ref Quad q)     => q.First ?? q.Second ?? q.Third ?? q.Fourth;
    [MethodImpl(MethodImplOptions.NoInlining)] public static object ReadMixed(ref Mixed m)   => m.Second == 0 ? m.First : null;

    [MethodImpl(MethodImplOptions.NoInlining)] public static object ProbePair()   { Pair p = default;   return ReadPair(ref p); }
    [MethodImpl(MethodImplOptions.NoInlining)] public static object ProbeTriple() { Triple t = default; return ReadTriple(ref t); }
    [MethodImpl(MethodImplOptions.NoInlining)] public static object ProbeQuad()   { Quad q = default;   return ReadQuad(ref q); }
    [MethodImpl(MethodImplOptions.NoInlining)] public static object ProbeMixed()  { Mixed m = default;  return ReadMixed(ref m); }

    public static int Main()
    {
        int acc = (ProbePair() is null ? 1 : 0) + (ProbeTriple() is null ? 1 : 0) +
                  (ProbeQuad() is null ? 1 : 0) + (ProbeMixed() is null ? 1 : 0);
        Console.WriteLine($"OK acc={acc}");
        return 100;
    }
}

Run with DOTNET_TieredCompilation=0, DOTNET_ReadyToRun=0, DOTNET_JitDisasm="ProbePair ProbeTriple ProbeQuad ProbeMixed" on windows-x64.

Current codegen

Repro:ProbePair() — prolog, 33 bytes total method size:

G_M000_IG01:
       sub      rsp, 56
       xor      eax, eax                      ; 2 bytes
       mov      qword ptr [rsp+0x28], rax     ; 5 bytes
       mov      qword ptr [rsp+0x30], rax     ; 5 bytes
                                              ; 12 bytes / 3 instructions

Repro:ProbeTriple() (24 bytes of slots) already takes the block-init path, so the 16-byte case is the odd one out.

Expected codegen

G_M000_IG01:
       sub      rsp, 56
       vxorps   xmm4, xmm4, xmm4                   ; 4 bytes
       vmovdqu  xmmword ptr [rsp+0x28], xmm4       ; 6 bytes
                                                   ; 10 bytes / 2 instructions

Method size 33 → 31 bytes. No frame-layout change, no frame growth, no new callee-save: xmm4 (Windows) / xmm8 (SysV) is the same volatile scratch the existing block-init path already uses. Unaligned vmovdqu and aligned vmovdqa are the same length, so an unaligned 16-byte range still wins.

The saving grows on large frames, where each scalar store pays disp32+SIB. Example from the corpus, Microsoft.CodeAnalysis.CSharp.Binder:BindSimpleProgramCompilationUnit(...) (FullOpts, 1312 → 1306 bytes):

base:  xor       eax, eax                           ; 2
       mov       qword ptr [rsp+0x90], rax          ; 8
       mov       qword ptr [rsp+0x98], rax          ; 8
diff:  vxorps    xmm4, xmm4, xmm4                   ; 4
       vmovdqa32 xmmword ptr [rsp+0x90], xmm4       ; 8

Impact

Target method: 33 → 31 bytes (−2), 3 → 2 prolog instructions, 2 → 1 store uop.

SuperPMI asmdiffs, windows-x64, Release JIT pair, all 12 default collections (JIT-EE GUID fbbaf45f-5b0e-4767-b962-5084c8caae77), base commit 4856f0c16f89d8cc625a62ee7bab7ed4afe14f9d:

Collection Delta bytes % Ctx w/ diffs Impr Regr
aspire.nativeaot −4,207 −0.03 1,127 1,127 0
aspnet2.run −2,142 −0.01 1,624 1,622 2
benchmarks.run −1,139 −0.01 654 652 2
benchmarks.run_pgo −3,381 −0.01 3,166 3,165 0
benchmarks.run_pgo_optrepeat −1,135 −0.01 653 651 2
coreclr_tests.run −30,557 −0.01 27,513 27,512 1
libraries.crossgen2 −8,901 −0.02 4,908 4,907 1
libraries.pmi −8,121 −0.01 4,727 4,721 6
libraries_tests.run −50,901 −0.01 47,579 47,575 3
libraries_tests_no_tiered_compilation.run −21,808 −0.01 11,953 11,949 4
realworld.run −1,349 −0.01 846 846 0
smoke_tests.nativeaot −1,979 −0.03 625 625 0
Total −135,620 −0.012% 105,375 105,352 21

Zero failing compiles in every collection, base and diff. No new asserts. Missing-context counts are identical in base and diff everywhere except libraries_tests.run, where the diff JIT has fewer (29 vs 38).

Known regressions. All 21 regressions are exactly +1 byte, and all have the same cause: when the integer zero register is already live for other reasons the xor is not actually saved, and on an rbp frame with disp8 the two 4-byte scalar stores (8 bytes) beat vxorps + vmovdqu (9 bytes). Representative case, Microsoft.Extensions.Logging.LoggerFactory:RefreshFilters(...) (Tier1, 279 → 280):

base:  xor     edi, edi                            ; edi is also the zero source for later uses
       mov     qword ptr [rbp-0x48], rdi           ; 4
       mov     qword ptr [rbp-0x40], rdi           ; 4
diff:  vxorps  xmm4, xmm4, xmm4                    ; 4   (xor edi,edi survives — still needed)
       vmovdqu xmmword ptr [rbp-0x48], xmm4        ; 5

Suppressing these would need initReg liveness information that is not available at this point in the prolog. 21 bytes against −135,641 bytes of improvements.

Throughput. superpmi.py tpdiff (PIN) on 4 of 12 collections before the session budget was exhausted: aspire.nativeaot +0.00%, aspnet2.run +0.01%, benchmarks.run +0.00%, benchmarks.run_pgo +0.01% of total instructions executed. Both JITs carry superpmi.py's "compiled with native PGO" warning, so sub-0.05% readings are within noise.

Measurement limitations:

  • PerfScore is unmeasured. Both JITs in this experiment were Release builds and RyuJIT only computes PerfScore under DEBUG/LATE_DISASM, so every collection reports Total PerfScore 0.0 / geomean 1.0. No execution-speed claim is made here; this is a code-size and instruction-count result.
  • tpdiff covers 4 of 12 collections.
  • Only windows-x64 collections were replayed. Linux/SysV AMD64 (different scratch register, different calleeRegArgMaskLiveIn, different frame offsets) was not diffed and is the one gap worth closing.
  • No runtime test suite was run against the prototype (JIT-only build). Correctness evidence is the repro's identical managed output (OK acc=4, exit 100 on both JITs) plus 105,375 replayed contexts with zero failing compiles.

Notes

  • Scope. AMD64 only in the prototype (#ifdef TARGET_AMD64). x86, arm, arm64, loongarch64, riscv64 and wasm are untouched; their thresholds and block-init helpers are unmodified. arm64 may want an analogous treatment via stp, but that is not evaluated here.
  • Threshold left alone. Lowering genCheckUseBlockInit's AMD64 genInitStkLclCnt > 4 to > 3 would be the obvious fix but is wrong: it would also admit sparse 4-slot sets whose [untrLclLo, untrLclHi) span covers unrelated locals, zeroing memory that the scalar path does not touch. The density check has to happen after offsets are assigned.
  • Same bytes zeroed. Requiring genInitStkLclCnt == 4 and untrLclHi - untrLclLo == 16 forces the range to be exactly 16 bytes fully covered by lvMustInit slots. A further guard requires that the scalar path would have written all 16 of those bytes — this excludes the TYP_STRUCT + compInitMem == false case where the scalar path writes only the GC slots (block init would then emit 10 bytes to replace a single 5-byte store). Nothing outside the must-init locals is touched, so no GC-reportable slot changes meaning.
  • initRegZeroed is already handled. genZeroInitFrameUsingBlockInit leaves *pInitRegZeroed untouched when it does not need the integer zero, so later prolog consumers (GS cookie, generic context arg) correctly re-materialize it. That is pre-existing behavior on the genUseBlockInit path; this only widens which methods take it.
  • Compiler::fgVarNeedsExplicitZeroInit is deliberately not changed. It consumes genCheckUseBlockInit's threshold to decide which locals can drop an explicit body-level zero init. Since this does not alter which locals are prolog-zeroed (lvMustInit is untouched) — only the instruction sequence used — that coupling stays sound. Revisiting it is a possible follow-up, not part of this.
  • AVX state. No new kind of instruction is introduced; the implicit zeroing of YMM4/ZMM4 and the AVX-state bookkeeping are identical to the existing block-init path, which already emits this sequence for larger ranges.
  • Related prior work: Use xmm for stack prolog zeroing rather than rep stos #32538 (merged) introduced the xmm prolog-zeroing path for the large end, replacing rep stosd; Use simd for small prolog zeroing (ia32/x64) #32442 was closed in its favor. This is the opposite end of the same threshold and does not overlap with either.

Prototype patch

Experimental patch
diff --git a/src/coreclr/jit/codegen.h b/src/coreclr/jit/codegen.h
index 2020048e2ac..616c1356f87 100644
--- a/src/coreclr/jit/codegen.h
+++ b/src/coreclr/jit/codegen.h
@@ -577,6 +577,9 @@ protected:
 
     void genZeroInitFrame(int untrLclHi, int untrLclLo, regNumber initReg, bool* pInitRegZeroed);
     void genZeroInitFrameUsingBlockInit(int untrLclHi, int untrLclLo, regNumber initReg, bool* pInitRegZeroed);
+#ifdef TARGET_AMD64
+    bool genCanZeroInitFrameUsingOneSimdStore(int untrLclHi, int untrLclLo);
+#endif
 
     void genReportGenericContextArg(regNumber initReg, bool* pInitRegZeroed);
 
diff --git a/src/coreclr/jit/codegencommon.cpp b/src/coreclr/jit/codegencommon.cpp
index 5b257d36dc9..c3af59312eb 100644
--- a/src/coreclr/jit/codegencommon.cpp
+++ b/src/coreclr/jit/codegencommon.cpp
@@ -4197,6 +4197,82 @@ regNumber CodeGen::genGetZeroReg(regNumber initReg, bool* pInitRegZeroed)
 #endif // !TARGET_ARM64
 }
 
+#ifdef TARGET_AMD64
+//-----------------------------------------------------------------------------
+// genCanZeroInitFrameUsingOneSimdStore: Determine whether the prolog zeroing work is a single
+//    dense 16-byte range that is cheaper to clear with one SIMD store.
+//
+// Arguments:
+//    untrLclHi - (Untracked locals High-Offset)  The upper bound offset (not inclusive) of the
+//                                               range that will be zeroed.
+//    untrLclLo - (Untracked locals Low-Offset)   The lower bound offset of that range.
+//
+// Return Value:
+//    True if genZeroInitFrameUsingBlockInit should be used even though genCheckUseBlockInit did
+//    not set genUseBlockInit.
+//
+// Notes:
+//    genCheckUseBlockInit runs before frame offsets are assigned, so it can only look at the
+//    number of int-sized slots to zero and has to stay conservative. When those slots turn out
+//    to form one contiguous 16-byte range that the scalar path would store in full, block init
+//    emits "xorps xmm, xmm" plus a single 16-byte store, which is one instruction and two bytes
+//    shorter than zeroing an integer register plus two pointer-sized stores.
+//
+//    Ranges the scalar path would only write partially (a struct with non-GC fields when
+//    compInitMem is false) are excluded, because block init would then be larger than the stores
+//    it replaces.
+//
+bool CodeGen::genCanZeroInitFrameUsingOneSimdStore(int untrLclHi, int untrLclLo)
+{
+    if (genUseBlockInit || (genInitStkLclCnt != (XMM_REGSIZE_BYTES / sizeof(int))) ||
+        ((untrLclHi - untrLclLo) != XMM_REGSIZE_BYTES))
+    {
+        return false;
+    }
+
+    // Tally the bytes the scalar path in genZeroInitFrame would store; this mirrors the loop there.
+    unsigned scalarZeroedBytes = 0;
+
+    for (unsigned varNum = 0; varNum < m_compiler->lvaCount; varNum++)
+    {
+        LclVarDsc* varDsc = m_compiler->lvaGetDesc(varNum);
+
+        if (!varDsc->lvMustInit || (varDsc->lvIsInReg() && !varDsc->IsLiveInOutOfHandler()) ||
+            m_compiler->lvaIsUnknownSizeLocal(varNum))
+        {
+            continue;
+        }
+
+        if (varDsc->TypeIs(TYP_STRUCT) && !m_compiler->info.compInitMem &&
+            (varDsc->lvExactSize() >= TARGET_POINTER_SIZE))
+        {
+            const unsigned slots  = (unsigned)m_compiler->lvaLclStackHomeSize(varNum) / REGSIZE_BYTES;
+            ClassLayout*   layout = varDsc->GetLayout();
+
+            for (unsigned i = 0; i < slots; i++)
+            {
+                scalarZeroedBytes += layout->IsGCPtr(i) ? REGSIZE_BYTES : 0;
+            }
+        }
+        else
+        {
+            scalarZeroedBytes += roundUp(m_compiler->lvaLclStackHomeSize(varNum), (unsigned)sizeof(int));
+        }
+    }
+
+    assert(regSet.tmpAllFree());
+    for (TempDsc* tempThis = regSet.tmpListBeg(); tempThis != nullptr; tempThis = regSet.tmpListNxt(tempThis))
+    {
+        if (varTypeIsGC(tempThis->tdTempType()))
+        {
+            scalarZeroedBytes += REGSIZE_BYTES;
+        }
+    }
+
+    return scalarZeroedBytes == XMM_REGSIZE_BYTES;
+}
+#endif // TARGET_AMD64
+
 //-----------------------------------------------------------------------------
 // genZeroInitFrame: Zero any untracked pointer locals and/or initialize memory for locspace
 //
@@ -4212,6 +4288,13 @@ void CodeGen::genZeroInitFrame(int untrLclHi, int untrLclLo, regNumber initReg,
 {
     assert(GetEmitter()->emitGeneratingPrologOrFuncletProlog());
 
+#ifdef TARGET_AMD64
+    if (genCanZeroInitFrameUsingOneSimdStore(untrLclHi, untrLclLo))
+    {
+        genUseBlockInit = true;
+    }
+#endif
+
     if (genUseBlockInit)
     {
         genZeroInitFrameUsingBlockInit(untrLclHi, untrLclLo, initReg, pInitRegZeroed);

Note

This issue was generated with GitHub Copilot.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIperformance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions