You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Lowering::LowerStoreIndir (src/coreclr/jit/lowerxarch.cpp) already retypes a setcc/compare producer to TYP_BYTE when it feeds a one-byte GT_STOREIND, which makes CodeGen::inst_SETCC skip the zero-extension (needsMovzx = !varTypeIsByte(type)). The isomorphic shape for locals — a byte-sized store to a local that is guaranteed to live in memory (lvDoNotEnregister) — gets no such treatment in Lowering::LowerStoreLoc, so the JIT emits setcc + movzx + a 1-byte store. Only the low byte of the producer's register is ever consumed by that store, so the movzx is dead work.
Minimal repro
usingSystem;usingSystem.Runtime.CompilerServices;staticclassProgram{publicstructState{publicboolReady;publicintValue;}publicstaticintSink;[MethodImpl(MethodImplOptions.NoInlining)]staticvoidObserve(refStatestate)=>Sink=state.Ready?1:0;// Target: 'state' is address-exposed, so state.Ready is a byte-sized store to a stack home.[MethodImpl(MethodImplOptions.NoInlining)]publicstaticvoidUpdateField(intx,inty){Statestate=default;Observe(refstate);state.Ready=x>y;Observe(refstate);}// Control: the STOREIND form is already optimized today.[MethodImpl(MethodImplOptions.NoInlining)]publicstaticvoidUpdateFieldIndir(refStatestate,intx,inty)=>state.Ready=x>y;// Negative control: the movzx is required here (bool return normalization).[MethodImpl(MethodImplOptions.NoInlining)]publicstaticboolUpdateReturn(intx,inty)=>x>y;staticintMain(){boolok=true;for(intx=-2;x<=2;x++){States=default;UpdateFieldIndir(refs,x,0);UpdateField(x,0);booldirect=x>0;if(s.Ready!=direct||UpdateReturn(x,0)!=direct||(Sink==1)!=direct){ok=false;Console.WriteLine("SETCC MISMATCH x="+x);}}Console.WriteLine(ok?"EQUIV-OK":"EQUIV-FAIL");return100;}}
Run with DOTNET_TieredCompilation=0 and DOTNET_JitDisasm=Update* (x64, FullOpts).
Current codegen
; Program:UpdateField(int,int) (FullOpts) -- Total bytes of code 59cmpebx,esi setg clmovzxrcx,cl ; redundant: only CL is storedmov byte ptr [rsp+0x20],cl; Program:UpdateFieldIndir(byref,int,int) (FullOpts) -- Total bytes of code 9cmpedx,r8d setg almov byte ptr [rcx],al ; STOREIND form: already has no movzx
Expected codegen
; Program:UpdateField(int,int) (FullOpts) -- Total bytes of code 56cmpebx,esi setg clmov byte ptr [rsp+0x20],cl
Program:UpdateReturn(int,int):bool is unchanged (setg al + movzx rax, al), so bool return normalization is not affected.
Largest single improvement: System.Reflection.Metadata.MetadataReader:InitializeTableReaders (Tier1), 16,620 → 16,548 bytes (−72); the textual diff is exactly 24 removed movzx rNN, rNNb instructions and nothing else. The next largest are −72/−71/−71/−71, all the same shape.
The reported +253 bytes of "regression" is confined to two Tier0 contexts in libraries_tests.run. Inspecting all 175 diffed .dasm pairs in that collection shows 173 methods shrink (−1,438 bytes total), 2 are byte-for-byte identical (Compare-Object finds no textual difference; 3,827→3,827 and 751→751) and zero methods grow; the delta comes from missing-context accounting asymmetry in that collection (base 37 / diff 29, which fluctuates run-to-run in standalone replays too), not from larger code.
Measurement limitations: SuperPMI reported PerfScore as unchanged (0.00%) for all 416 diffed contexts even though each removes one or more 1-µop movzx; the per-method measurement shows a small improvement, so in either reading there is no PerfScore regression. tpdiff was not run (PIN unavailable). No microbenchmark — the claim is static code size, not an end-to-end speedup.
Notes
Scope: x64/x86 (lowerxarch.cpp). ARM64 is untouched; hoisting the predicate into the shared LowerStoreLocCommon so ARM64 can reuse it is a possible follow-up.
Correctness assumptions: SETcc writes an 8-bit register and the consumer is a 1-byte store, so the bits the movzx clears are never read. lvDoNotEnregister guarantees the destination is a stack home, which keeps the transform away from enregistered small locals where lvNormalizeOnStore requires a normalized register value. Loads from such locals are emitted from the local's own small type (ins_Load(TYP_UBYTE)), so no consumer observes slot padding. LIR single-use means storeLoc->Data() has no other consumer, the same reasoning LowerStoreIndir already relies on. No LSRA change is needed: source register candidates derive from the store node's type, which is unchanged.
Remaining risks: frequency is low (0.015% of contexts); APX/EnableApxZU interaction is reasoned about but not measured (no APX collection in the default set — note that with ZU, inst_SETCC already skips the movzx, so the change makes that branch unreachable for these sites rather than requesting the longer encoding); x86 (32-bit) was not replayed; a tighter IsAddressExposed() predicate is an alternative to lvDoNotEnregister.
Prototype patch
Experimental patch
diff --git a/src/coreclr/jit/lowerxarch.cpp b/src/coreclr/jit/lowerxarch.cpp
index 7fbed52a6d1..d7ece7fc20d 100644
--- a/src/coreclr/jit/lowerxarch.cpp+++ b/src/coreclr/jit/lowerxarch.cpp@@ -66,6 +66,14 @@ GenTree* Lowering::LowerStoreLoc(GenTreeLclVarCommon* storeLoc)
verifyLclFldDoNotEnregister(storeLoc->GetLclNum());
}
+ // Optimization: do not unnecessarily zero-extend the result of setcc when storing it to a one-byte+ // stack location - only the low byte of the producer's register is consumed by the store.+ if (varTypeIsByte(storeLoc) && m_compiler->lvaGetDesc(storeLoc)->lvDoNotEnregister &&+ (storeLoc->Data()->OperIsCompare() || storeLoc->Data()->OperIs(GT_SETCC)))+ {+ storeLoc->Data()->ChangeType(TYP_BYTE);+ }+
ContainCheckStoreLoc(storeLoc);
return storeLoc->gtNext;
}
Lowering::LowerStoreIndir(src/coreclr/jit/lowerxarch.cpp) already retypes asetcc/compare producer toTYP_BYTEwhen it feeds a one-byteGT_STOREIND, which makesCodeGen::inst_SETCCskip the zero-extension (needsMovzx = !varTypeIsByte(type)). The isomorphic shape for locals — a byte-sized store to a local that is guaranteed to live in memory (lvDoNotEnregister) — gets no such treatment inLowering::LowerStoreLoc, so the JIT emitssetcc+movzx+ a 1-byte store. Only the low byte of the producer's register is ever consumed by that store, so themovzxis dead work.Minimal repro
Run with
DOTNET_TieredCompilation=0andDOTNET_JitDisasm=Update*(x64, FullOpts).Current codegen
Expected codegen
Program:UpdateReturn(int,int):boolis unchanged (setg al+movzx rax, al), soboolreturn normalization is not affected.Impact
Target method: 59 → 56 bytes, 20 → 19 instructions, PerfScore 15.50 → 15.25 (Checked-JIT footers). Both runtimes print
EQUIV-OKand exit 100.SuperPMI
asmdiffsagainst a prototype (below), x64 Release, default 12-collection set, 2,787,663 contexts (1,013,313 MinOpts / 1,774,350 FullOpts):All replays cleanin every collectionPer-collection:
aspire.nativeaot −34,aspnet2.run −110,benchmarks.run −80,benchmarks.run_pgo −69,benchmarks.run_pgo_optrepeat −80,coreclr_tests.run −136,libraries.crossgen2 −159,libraries.pmi −117,libraries_tests.run −1,185,libraries_tests_no_tiered_compilation −356,realworld.run −83,smoke_tests.nativeaot 0.Largest single improvement:
System.Reflection.Metadata.MetadataReader:InitializeTableReaders(Tier1), 16,620 → 16,548 bytes (−72); the textual diff is exactly 24 removedmovzx rNN, rNNbinstructions and nothing else. The next largest are −72/−71/−71/−71, all the same shape.The reported +253 bytes of "regression" is confined to two Tier0 contexts in
libraries_tests.run. Inspecting all 175 diffed.dasmpairs in that collection shows 173 methods shrink (−1,438 bytes total), 2 are byte-for-byte identical (Compare-Objectfinds no textual difference; 3,827→3,827 and 751→751) and zero methods grow; the delta comes from missing-context accounting asymmetry in that collection (base 37 / diff 29, which fluctuates run-to-run in standalone replays too), not from larger code.Measurement limitations: SuperPMI reported PerfScore as unchanged (0.00%) for all 416 diffed contexts even though each removes one or more 1-µop
movzx; the per-method measurement shows a small improvement, so in either reading there is no PerfScore regression.tpdiffwas not run (PIN unavailable). No microbenchmark — the claim is static code size, not an end-to-end speedup.Notes
lowerxarch.cpp). ARM64 is untouched; hoisting the predicate into the sharedLowerStoreLocCommonso ARM64 can reuse it is a possible follow-up.SETccwrites an 8-bit register and the consumer is a 1-byte store, so the bits themovzxclears are never read.lvDoNotEnregisterguarantees the destination is a stack home, which keeps the transform away from enregistered small locals wherelvNormalizeOnStorerequires a normalized register value. Loads from such locals are emitted from the local's own small type (ins_Load(TYP_UBYTE)), so no consumer observes slot padding. LIR single-use meansstoreLoc->Data()has no other consumer, the same reasoningLowerStoreIndiralready relies on. No LSRA change is needed: source register candidates derive from the store node's type, which is unchanged.EnableApxZUinteraction is reasoned about but not measured (no APX collection in the default set — note that with ZU,inst_SETCCalready skips themovzx, so the change makes that branch unreachable for these sites rather than requesting the longer encoding); x86 (32-bit) was not replayed; a tighterIsAddressExposed()predicate is an alternative tolvDoNotEnregister.Prototype patch
Experimental patch
Note
This issue was generated with GitHub Copilot.