fix(compile): make autotuning and parallel kernel compilation work on flagos - #59
Open
lvyufeng wants to merge 1 commit into
Open
fix(compile): make autotuning and parallel kernel compilation work on flagos#59lvyufeng wants to merge 1 commit into
lvyufeng wants to merge 1 commit into
Conversation
… flagos The torch.compile integration merged in flagos-ai#41 only ever compiled single-Linear models in its tests, which need neither autotuning nor more than one Triton kernel. Two independent failures hid behind that. Both reproduce on any graph with a couple of stacked Linears or a LayerNorm. Autotuning needs a constructible event. InductorBenchmarker.get_event_pairs times candidate configs with torch.cuda.Event(enable_timing=True). In the CPU-only wheel this build pairs with an external libtorch_cuda.so, that binding was never compiled, so torch.cuda substitutes a placeholder from torch._utils._dummy_type whose __new__ raises "Tried to instantiate dummy base class Event". flagos.Event subclassed it and inherited the failure. flagos.Event now picks its base class by lineage: on a vendor torch build it still subclasses torch.cuda.Event, and when that is a dummy it subclasses the device-agnostic torch.Event, which dispatches record/block/query/ elapsedTime to c10::flagos::DeviceGuardImpl (csrc/runtime/guard.h). Timing stays a real device measurement, and since every vendor under csrc/runtime/accelerator/ implements that ABI, the fallback is portable rather than NVIDIA-specific. Note the fix has to land here: patching triton.testing.do_bench does not help, because inductor reaches the benchmarker through triton_heuristics.benchmark_all_configs -> bench -> benchmarker.benchmark_gpu, not through do_bench. Compile workers need torch_fl. Inductor's default worker_start_method, "subprocess", starts workers as a bare `sys.executable -m torch._inductor.compile_worker` that imports only torch and triton. flagos lives behind PrivateUse1, so such a worker has no accelerator: triton's CudaDriver.is_active() asks torch.cuda.is_available(), gets False, and the worker dies with "Could not find an active GPU backend". "fork" inherits this process, torch_fl included, so workers come up already seeing the device -- and compilation stays parallel, unlike compile_threads = 1 (Qwen3-0.6B: 31.9s forked vs 40.8s serial). Both overrides are scoped to this build by probing for a missing torch._C CUDA binding, so a vendor torch install keeps inductor's defaults. tests/integration/test_compile_autotune.py guards both: stacked Linears, normalizations, reductions, multi-kernel backward, dynamic shapes and max-autotune, plus a direct check that the autotuner's own Event call works. On a cleared TORCHINDUCTOR_CACHE_DIR, 7 of its 8 tests fail before these changes. CI ran no compile tests at all, so .github/configs/cuda.yml now runs both compile files -- in one pytest invocation on purpose. The worker pool is created lazily and shared, so a file run on its own can be served entirely before the pool spins up, which is exactly how the worker failure stayed hidden until a multi-file run reproduced it. Measured on one A100, fp32, compiled vs eager on flagos: Qwen3-0.6B forward 2.24x (35.6ms -> 15.9ms, numerics matching eager at rtol/atol 2e-2), elementwise chain 4096x4096 9.18x, transformer block 1.41x, matmul-bound MLP 1.06x, and 0.92x at 64x512 where launch overhead exceeds the saving. tests/perf/bench_compile.py had never been run and could not be: it called torch.gelu (nonexistent), read torch.os.environ, recognised only the "privateuseone" spelling of the device, and imported torch before torch_fl -- which the docs now state as a hard requirement, since torch_fl preloads the libtorch_cuda.so that torch.cuda depends on. The test suite gets away without it because conftest imports torch_fl during collection. Also documents a third bug found while benchmarking and left unfixed: convolutions do not compile. Inductor prefers channels_last for conv on GPU, and while the flagos conv kernel honours that layout, its fake/meta kernel still predicts contiguous strides, so inductor rejects the graph on a stride mismatch. Eager never hits it, since it is the layout pass that produces a channels_last input. Reproduce with bench_compile.py --model=conv.
lvyufeng
force-pushed
the
fix/compile-autotune-and-workers
branch
from
August 6, 2026 02:58
9dd33f4 to
e375d3e
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The torch.compile integration merged in #41 only ever compiled single-
Linearmodels in its tests, which need neither autotuning nor more than one Triton
kernel. Two independent failures hid behind that. Both reproduce on any graph
with a couple of stacked Linears or a LayerNorm.
Autotuning needs a constructible event
InductorBenchmarker.get_event_pairstimes candidate configs withtorch.cuda.Event(enable_timing=True). In the CPU-only wheel this build pairswith an external
libtorch_cuda.so, that binding was never compiled, sotorch.cudasubstitutes a placeholder fromtorch._utils._dummy_typewhose__new__raises "Tried to instantiate dummy base class Event".flagos.Eventsubclassed it and inherited the failure.flagos.Eventnow picks its base class by lineage: on a vendor torch build itstill subclasses
torch.cuda.Event, and when that is a dummy it subclasses thedevice-agnostic
torch.Event, which dispatches record/block/query/elapsedTimeto
c10::flagos::DeviceGuardImpl(csrc/runtime/guard.h). Timing stays a realdevice measurement, and since every vendor under
csrc/runtime/accelerator/implements that ABI, the fallback is portable rather than NVIDIA-specific.
The fix has to land here. Patching
triton.testing.do_benchdoes not help,because inductor reaches the benchmarker through
triton_heuristics.benchmark_all_configs -> bench -> benchmarker.benchmark_gpu,not through
do_bench.Compile workers need torch_fl
Inductor's default
worker_start_method,"subprocess", starts workers as abare
sys.executable -m torch._inductor.compile_workerthat imports only torchand triton. flagos lives behind PrivateUse1, so such a worker has no
accelerator: triton's
CudaDriver.is_active()askstorch.cuda.is_available(),gets
False, and the worker dies with "Could not find an active GPU backend"."fork"inherits this process,torch_flincluded, so workers come up alreadyseeing the device — and compilation stays parallel, unlike
compile_threads = 1(Qwen3-0.6B: 31.9s forked vs 40.8s serial). Both overrides are scoped to this
build by probing for a missing
torch._CCUDA binding, so a vendor torchinstall keeps inductor's defaults.
Testing
tests/integration/test_compile_autotune.pyguards both: stacked Linears,normalizations, reductions, multi-kernel backward, dynamic shapes and
max-autotune, plus a direct check that the autotuner's own
Eventcall works.On a cleared
TORCHINDUCTOR_CACHE_DIR, 7 of its 8 tests fail before thesechanges.
CI ran no compile tests at all, so
.github/configs/cuda.ymlnow runs bothcompile files — in one pytest invocation on purpose. The worker pool is created
lazily and shared, so a file run on its own can be served entirely before the
pool spins up, which is exactly how the worker failure stayed hidden until a
multi-file run reproduced it.
Verified on one A100 (torch 2.10 CPU wheel + external cu128
libtorch_cuda.so):tests/integration/ops -m "main_ops and not flaggems_python and not flaggems_cpp": 117 passed, 15 skipped, 1 xfailedruff check/ruff format --checkclean on all touched filesDeprecationWarningstress-tested over 6 sequentialcompile/pool cycles, no hangs;
spawnwas also tried and fails (1 of 4)Performance
Measured fp32, compiled vs eager on flagos:
Qwen3 numerics match eager at rtol/atol 2e-2. The pattern is what inductor's
fusion predicts: the win comes from collapsing elementwise chains, matmul-bound
graphs still call cuBLAS, and at small sizes launch overhead exceeds the saving.
Also in here
tests/perf/bench_compile.pyhad never been run and could not be: it calledtorch.gelu(nonexistent), readtorch.os.environ, recognised only the"privateuseone"spelling of the device, and imported torch before torch_fl —which the docs now state as a hard requirement, since torch_fl preloads the
libtorch_cuda.sothattorch.cudadepends on. The test suite gets awaywithout ordering its imports because conftest imports torch_fl during
collection.
Known limitation documented, not fixed
Convolutions do not compile. Inductor prefers
channels_lastfor conv on GPU,and while the flagos conv kernel honours that layout, its fake/meta kernel
still predicts contiguous strides, so inductor rejects the graph on a stride
mismatch in
aten.convolution.default. Confirmed by direct comparison on a16x3x224x224
channels_lastinput — flagos real(3211264, 1, 14336, 64)vsfake
(3211264, 50176, 224, 1), while cuda agrees with itself. Eager neverhits it, since it is the layout pass that produces a
channels_lastinput.Reproduce with
bench_compile.py --model=conv. Filed as Limitations #6.