Skip to content

Add a low-cost tracing subsystem - #836

Open
vchuravy wants to merge 3 commits into
mainfrom
vc/tracing-v2
Open

vchuravy wants to merge 3 commits into
mainfrom
vc/tracing-v2

Conversation

@vchuravy

@vchuravy vchuravy commented Oct 4, 2026

Copy link
Copy Markdown
Member

Adds a low-cost tracing subsystem that puts named ranges on the timeline of a tracing profiler. This supersedes #703 and #66.

@profiling_range "volume integral" domain = "Trixi" begin
    volume_integral!(du, u, backend)
end
profiling_mark("converged")

Kernel launches are also recorded as ranges, named after the kernel.

Design

NVTX, ITT and roctx annotate host threads, and the profiler attributes the device work launched within a range to that range. So which profiler records a range depends on what the process runs under, not on the backend. For example, running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend under VTune gives ITT tasks. For this reason the API takes no backend argument, and there are no KernelInterface hooks. Every range goes to all registered KernelAbstractions.Tracers.

  • Cost when no profiler is attached: a single atomic load (about 0.6 ns), because the tracer list is copy-on-write. The label isn't evaluated, so it can use string interpolation for free. The tracing part of a kernel launch is kept in a separate non-inlined function.
  • Ranges use start/end, not push/pop, so they may end on another thread and needn't nest. A range is ended by the tracers that were registered when it started.
  • @profiling_range uses a scope-free :tryfinally, like @time. The range is ended if the expression throws, and assignments inside it remain visible afterwards.

Tracers

Each tracer registers itself only when its profiler is attached:

Tracer Trigger Active when
NVTXExt NVTX.jl (also loaded by CUDA.jl) NVTX.isactive(), i.e. under nsys
ROCTXExt AMDGPU.jl under rocprofv3 (ROCP_TOOL_LIBRARIES) or legacy rocprof (HSA_TOOLS_LIB)
IntelITTExt IntelITT.jl IntelITT.isactive(), i.e. under VTune
NVTXTTracer built in JULIA_KA_NVTXT=1 or =path-%p.nvtxt

NVTXTTracer revives #66. It writes the NVTXT text format, which Nsight Systems can import with ImportNvtxt, so you can trace without a profiler attached.

Testing

  • The test suite covers the macro, registration, ranges, markers, ranges across many tasks, multiple tracers, NVTXT output (including the environment variable, in a subprocess), and the backend testsuite with a tracer registered.
  • Verified by hand under nsys profile --trace=nvtx: nvtx_sum shows Demo:step 1 and KernelAbstractions:scale! ranges.
  • ROCTX is not covered by CI. No AMDGPU release is compatible with KA 0.10 yet. I exercised the extension's code against /opt/rocm/lib/libroctx64.so (ROCm 7.2) using a stand-in AMDGPU module. I could not test it under rocprofv3, which isn't installed.

🤖 Generated with Claude Code

`@profiling_range` and `profiling_mark` put named ranges and markers on
the timeline of a tracing profiler, and kernel launches are annotated
with the kernel's name. Profilers like NVTX, ITT and roctx annotate host
threads and correlate device work themselves, so ranges go to every
registered `Tracer` regardless of backend. With none registered, an
annotation costs one atomic load.

Tracers: NVTX (NVTXExt), roctx (ROCTXExt, via AMDGPU), Intel ITT
(IntelITTExt), each registering only under its profiler, and a built-in
NVTXT writer enabled with `JULIA_KA_NVTXT`.

Supersedes #703 and #66.

Assisted-by: Claude Code (Opus 5.5)
Comment on lines +3 to +6
KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as
NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a
stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a
single atomic load, and the label isn't even built.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as
NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a
stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a
single atomic load, and the label isn't even built.
KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as
NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a
stretch of kernels belongs to. Annotations are cheap when no profiler is listening.

Comment on lines +33 to +37
Ranges are recorded on the host threads of the process, which is how NVTX, ITT and
roctx work: it is the profiler that attributes the device work launched within a range to
it. So which profiler records the ranges depends on what the process runs under, not on the
backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend
under VTune gives ITT tasks.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Ranges are recorded on the host threads of the process, which is how NVTX, ITT and
roctx work: it is the profiler that attributes the device work launched within a range to
it. So which profiler records the ranges depends on what the process runs under, not on the
backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend
under VTune gives ITT tasks.
Ranges are recorded on the host threads of the process, which is how NVTX, ITT and
roctx work: it is the profiler that attributes the device work launched within a range to
it.

This was referenced Oct 4, 2026
@codecov

codecov Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 0% with 234 lines in your changes missing coverage. Please review.
✅ Project coverage is 0.00%. Comparing base (e70abc3) to head (36ef5dd).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/profiler.jl 0.00% 100 Missing ⚠️
src/profiling.jl 0.00% 68 Missing ⚠️
ext/ROCTXExt.jl 0.00% 29 Missing ⚠️
ext/IntelITTExt.jl 0.00% 12 Missing ⚠️
ext/NVTXExt.jl 0.00% 11 Missing ⚠️
src/backend_launch.jl 0.00% 11 Missing ⚠️
src/KernelAbstractions.jl 0.00% 3 Missing ⚠️

❗ There is a different number of reports uploaded between BASE (e70abc3) and HEAD (36ef5dd). Click for more details.

HEAD has 24 uploads less than BASE
Flag BASE (e70abc3) HEAD (36ef5dd)
44 20
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #836       +/-   ##
==========================================
- Coverage   79.06%   0.00%   -79.07%     
==========================================
  Files          24      27        +3     
  Lines        2040    2075       +35     
==========================================
- Hits         1613       0     -1613     
- Misses        427    2075     +1648     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Records the ranges, markers and kernel launches of an expression with a
temporary tracer, and summarizes them per name or, with `trace = true`,
lists them in order. By default kernel launches synchronize their
backend while profiling, so that their ranges measure execution rather
than launch; tracers opt into this with `synchronizes_launches`.

Assisted-by: Claude Code (Opus 5.5)
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main 580a8e1... main / 580a8e1...
const/@Const/Float32/262144 0.526 ± 0.02 ms 0.53 ± 0.017 ms 0.993 ± 0.05
const/@Const/Float32/65536 0.175 ± 0.0093 ms 0.171 ± 0.0085 ms 1.02 ± 0.074
const/@Const/Float64/262144 0.904 ± 0.02 ms 0.894 ± 0.02 ms 1.01 ± 0.032
const/@Const/Float64/65536 0.369 ± 0.018 ms 0.368 ± 0.014 ms 1 ± 0.062
const/unmarked/Float32/262144 2.27 ± 0.026 ms 2.27 ± 0.013 ms 1 ± 0.013
const/unmarked/Float32/65536 0.598 ± 0.014 ms 0.599 ± 0.016 ms 0.999 ± 0.035
const/unmarked/Float64/262144 3.13 ± 0.026 ms 3.11 ± 0.041 ms 1.01 ± 0.016
const/unmarked/Float64/65536 0.811 ± 0.019 ms 0.817 ± 0.016 ms 0.992 ± 0.03
launch/3D static workgroup, dynamic ndrange 10.7 ± 1 μs 11 ± 0.45 μs 0.975 ± 0.1
launch/3D static workgroup, static ndrange 11.7 ± 0.68 μs 10.7 ± 0.56 μs 1.09 ± 0.085
launch/dynamic workgroup, dynamic ndrange 11.9 ± 0.46 μs 12.1 ± 0.55 μs 0.978 ± 0.059
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 12 ± 0.67 μs 12.1 ± 40 μs 0.991 ± 3.2
launch/static workgroup, dynamic ndrange 10.9 ± 0.7 μs 11 ± 0.63 μs 0.99 ± 0.086
launch/static workgroup, static ndrange 10.6 ± 0.49 μs 10.8 ± 0.61 μs 0.977 ± 0.071
partition/dynamic workgroup, dynamic ndrange 0.058 ± 0.0011 μs 0.0586 ± 0.00083 μs 0.989 ± 0.024
partition/static workgroup, dynamic ndrange 0.0602 ± 0.011 μs 0.0549 ± 0.012 μs 1.1 ± 0.31
partition/static workgroup, static ndrange 2.48 ± 0.001 ns 1.55 ± 0.01 ns 1.59 ± 0.01
saxpy/default/Float16/1024 0.0559 ± 0.0044 ms 0.0545 ± 0.0044 ms 1.03 ± 0.12
saxpy/default/Float16/1048576 1.47 ± 0.025 ms 1.47 ± 0.024 ms 0.998 ± 0.024
saxpy/default/Float16/16384 0.0701 ± 0.0098 ms 0.0735 ± 0.0081 ms 0.953 ± 0.17
saxpy/default/Float16/2048 0.0563 ± 0.0036 ms 0.0579 ± 0.0042 ms 0.972 ± 0.094
saxpy/default/Float16/256 11.3 ± 0.36 μs 11.7 ± 41 μs 0.97 ± 3.4
saxpy/default/Float16/262144 0.409 ± 0.018 ms 0.41 ± 0.019 ms 0.998 ± 0.064
saxpy/default/Float16/32768 0.0952 ± 0.0072 ms 0.0931 ± 0.0065 ms 1.02 ± 0.11
saxpy/default/Float16/4096 0.0632 ± 0.0068 ms 0.0597 ± 0.005 ms 1.06 ± 0.14
saxpy/default/Float16/512 0.0535 ± 0.0034 ms 11.9 ± 41 μs 4.49 ± 16
saxpy/default/Float16/64 11.3 ± 0.71 μs 11.7 ± 41 μs 0.965 ± 3.4
saxpy/default/Float16/65536 0.138 ± 0.0082 ms 0.14 ± 0.01 ms 0.989 ± 0.093
saxpy/default/Float32/1024 0.0524 ± 0.036 ms 0.0537 ± 0.0037 ms 0.976 ± 0.67
saxpy/default/Float32/1048576 0.765 ± 0.075 ms 0.757 ± 0.055 ms 1.01 ± 0.12
saxpy/default/Float32/16384 0.0592 ± 0.0065 ms 0.0615 ± 0.0065 ms 0.963 ± 0.15
saxpy/default/Float32/2048 0.0552 ± 0.0045 ms 0.0556 ± 0.0034 ms 0.993 ± 0.1
saxpy/default/Float32/256 11.3 ± 1.1 μs 11.5 ± 8.9 μs 0.981 ± 0.77
saxpy/default/Float32/262144 0.213 ± 0.027 ms 0.214 ± 0.028 ms 0.996 ± 0.18
saxpy/default/Float32/32768 0.0708 ± 0.0068 ms 0.072 ± 0.0069 ms 0.984 ± 0.13
saxpy/default/Float32/4096 0.0562 ± 0.0034 ms 0.0568 ± 0.0072 ms 0.99 ± 0.14
saxpy/default/Float32/512 14.5 ± 5.9 μs 14.8 ± 41 μs 0.978 ± 2.8
saxpy/default/Float32/64 11.2 ± 0.47 μs 11.4 ± 39 μs 0.983 ± 3.4
saxpy/default/Float32/65536 0.0932 ± 0.011 ms 0.092 ± 0.012 ms 1.01 ± 0.17
saxpy/default/Float64/1024 0.0546 ± 0.011 ms 0.0556 ± 0.007 ms 0.983 ± 0.23
saxpy/default/Float64/1048576 1.22 ± 0.14 ms 1.25 ± 0.14 ms 0.973 ± 0.16
saxpy/default/Float64/16384 0.0649 ± 0.0074 ms 0.0643 ± 0.0079 ms 1.01 ± 0.17
saxpy/default/Float64/2048 0.0571 ± 0.0035 ms 0.0554 ± 0.0049 ms 1.03 ± 0.11
saxpy/default/Float64/256 14.1 ± 4.9 μs 14.8 ± 43 μs 0.955 ± 2.8
saxpy/default/Float64/262144 0.272 ± 0.034 ms 0.276 ± 0.04 ms 0.983 ± 0.19
saxpy/default/Float64/32768 0.076 ± 0.0073 ms 0.0791 ± 0.011 ms 0.96 ± 0.16
saxpy/default/Float64/4096 0.0548 ± 0.009 ms 0.0533 ± 0.0083 ms 1.03 ± 0.23
saxpy/default/Float64/512 14.7 ± 1.9 μs 0.0519 ± 0.037 ms 0.283 ± 0.21
saxpy/default/Float64/64 0.0534 ± 0.043 ms 11.1 ± 0.26 μs 4.8 ± 3.9
saxpy/default/Float64/65536 0.104 ± 0.01 ms 0.106 ± 0.014 ms 0.985 ± 0.16
saxpy/static workgroup=(1024,)/Float16/1024 0.0543 ± 0.0032 ms 0.0547 ± 0.0026 ms 0.993 ± 0.075
saxpy/static workgroup=(1024,)/Float16/1048576 1.48 ± 0.024 ms 1.48 ± 0.019 ms 0.997 ± 0.021
saxpy/static workgroup=(1024,)/Float16/16384 0.0724 ± 0.0082 ms 0.0738 ± 0.0073 ms 0.982 ± 0.15
saxpy/static workgroup=(1024,)/Float16/2048 0.0585 ± 0.0033 ms 0.0564 ± 0.0048 ms 1.04 ± 0.11
saxpy/static workgroup=(1024,)/Float16/256 0.0501 ± 0.042 ms 11.9 ± 0.38 μs 4.22 ± 3.6
saxpy/static workgroup=(1024,)/Float16/262144 0.409 ± 0.02 ms 0.41 ± 0.02 ms 0.998 ± 0.068
saxpy/static workgroup=(1024,)/Float16/32768 0.0959 ± 0.0077 ms 0.0954 ± 0.0063 ms 1.01 ± 0.1
saxpy/static workgroup=(1024,)/Float16/4096 0.0609 ± 0.0055 ms 0.0585 ± 0.0046 ms 1.04 ± 0.12
saxpy/static workgroup=(1024,)/Float16/512 11.7 ± 38 μs 0.0555 ± 0.0031 ms 0.21 ± 0.69
saxpy/static workgroup=(1024,)/Float16/64 11.5 ± 0.21 μs 11.7 ± 10 μs 0.982 ± 0.85
saxpy/static workgroup=(1024,)/Float16/65536 0.14 ± 0.0076 ms 0.14 ± 0.0084 ms 1 ± 0.081
saxpy/static workgroup=(1024,)/Float32/1024 0.0524 ± 0.036 ms 0.0552 ± 0.0079 ms 0.949 ± 0.67
saxpy/static workgroup=(1024,)/Float32/1048576 0.8 ± 0.064 ms 0.774 ± 0.066 ms 1.03 ± 0.12
saxpy/static workgroup=(1024,)/Float32/16384 0.0599 ± 0.0065 ms 0.059 ± 0.007 ms 1.01 ± 0.16
saxpy/static workgroup=(1024,)/Float32/2048 0.0556 ± 0.0066 ms 0.0567 ± 0.0034 ms 0.979 ± 0.13
saxpy/static workgroup=(1024,)/Float32/256 11.8 ± 41 μs 0.0537 ± 0.036 ms 0.219 ± 0.78
saxpy/static workgroup=(1024,)/Float32/262144 0.221 ± 0.022 ms 0.219 ± 0.027 ms 1.01 ± 0.16
saxpy/static workgroup=(1024,)/Float32/32768 0.0722 ± 0.0067 ms 0.0721 ± 0.008 ms 1 ± 0.14
saxpy/static workgroup=(1024,)/Float32/4096 0.052 ± 0.0078 ms 0.0508 ± 0.0089 ms 1.02 ± 0.24
saxpy/static workgroup=(1024,)/Float32/512 18.7 ± 43 μs 0.0495 ± 0.044 ms 0.379 ± 0.94
saxpy/static workgroup=(1024,)/Float32/64 11.8 ± 38 μs 12.4 ± 42 μs 0.95 ± 4.5
saxpy/static workgroup=(1024,)/Float32/65536 0.0963 ± 0.0094 ms 0.0943 ± 0.011 ms 1.02 ± 0.15
saxpy/static workgroup=(1024,)/Float64/1024 0.0553 ± 0.01 ms 0.0554 ± 0.0036 ms 0.997 ± 0.2
saxpy/static workgroup=(1024,)/Float64/1048576 1.28 ± 0.23 ms 1.27 ± 0.16 ms 1.01 ± 0.22
saxpy/static workgroup=(1024,)/Float64/16384 0.0657 ± 0.0064 ms 0.0624 ± 0.0056 ms 1.05 ± 0.14
saxpy/static workgroup=(1024,)/Float64/2048 0.0567 ± 0.0044 ms 0.0556 ± 0.0054 ms 1.02 ± 0.13
saxpy/static workgroup=(1024,)/Float64/256 0.0554 ± 0.043 ms 0.0551 ± 0.0031 ms 1.01 ± 0.78
saxpy/static workgroup=(1024,)/Float64/262144 0.281 ± 0.035 ms 0.27 ± 0.034 ms 1.04 ± 0.18
saxpy/static workgroup=(1024,)/Float64/32768 0.0792 ± 0.0083 ms 0.0802 ± 0.01 ms 0.987 ± 0.16
saxpy/static workgroup=(1024,)/Float64/4096 0.054 ± 0.0089 ms 0.055 ± 0.008 ms 0.983 ± 0.22
saxpy/static workgroup=(1024,)/Float64/512 0.0533 ± 0.036 ms 0.0548 ± 0.0057 ms 0.973 ± 0.67
saxpy/static workgroup=(1024,)/Float64/64 12.4 ± 43 μs 0.0503 ± 0.041 ms 0.247 ± 0.88
saxpy/static workgroup=(1024,)/Float64/65536 0.11 ± 0.012 ms 0.105 ± 0.013 ms 1.04 ± 0.17
time_to_load 0.494 ± 0.0071 s 0.491 ± 0.0071 s 1.01 ± 0.02
main 580a8e1... main / 580a8e1...
const/@Const/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/@Const/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float32/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/262144 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
const/unmarked/Float64/65536 9 allocs: 0.203 kB 9 allocs: 0.203 kB 1
launch/3D static workgroup, dynamic ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/3D static workgroup, static ndrange 9 allocs: 0.219 kB 9 allocs: 0.219 kB 1
launch/dynamic workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, dynamic ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
launch/static workgroup, static ndrange 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/16384 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float32/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float32/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/default/Float64/32768 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/default/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/default/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/default/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/16384 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/static workgroup=(1024,)/Float16/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float16/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float16/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float32/32768 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float32/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float32/65536 8 allocs: 0.141 kB 12 allocs: 0.25 kB 0.562
saxpy/static workgroup=(1024,)/Float64/1024 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/16384 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/256 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/262144 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
saxpy/static workgroup=(1024,)/Float64/32768 12 allocs: 0.25 kB 8 allocs: 0.141 kB 1.78
saxpy/static workgroup=(1024,)/Float64/4096 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/512 8 allocs: 0.141 kB 8 allocs: 0.141 kB 1
saxpy/static workgroup=(1024,)/Float64/64 5 allocs: 0.0938 kB 5 allocs: 0.0938 kB 1
saxpy/static workgroup=(1024,)/Float64/65536 12 allocs: 0.25 kB 12 allocs: 0.25 kB 1
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

@vchuravy
vchuravy added this pull request to stack #838 October 4, 2026 08:38
Tracers are global, so `@profile` recorded every task in the process,
including other profiles. It now records only the task running its
expression and the tasks spawned from it, tracked with a ScopedValue,
numbers those tasks, and nests the trace per task rather than per
thread. It warns about ranges still open when it finishes, i.e. tasks
it wasn't waited for.

Assisted-by: Claude Code (Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant