Conversation
`@profiling_range` and `profiling_mark` put named ranges and markers on the timeline of a tracing profiler, and kernel launches are annotated with the kernel's name. Profilers like NVTX, ITT and roctx annotate host threads and correlate device work themselves, so ranges go to every registered `Tracer` regardless of backend. With none registered, an annotation costs one atomic load. Tracers: NVTX (NVTXExt), roctx (ROCTXExt, via AMDGPU), Intel ITT (IntelITTExt), each registering only under its profiler, and a built-in NVTXT writer enabled with `JULIA_KA_NVTXT`. Supersedes #703 and #66. Assisted-by: Claude Code (Opus 5.5)
vchuravy
commented
Oct 4, 2026
Comment on lines
+3
to
+6
| KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as | ||
| NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a | ||
| stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a | ||
| single atomic load, and the label isn't even built. |
Member
Author
There was a problem hiding this comment.
Suggested change
| KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as | |
| NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a | |
| stretch of kernels belongs to. Annotations are cheap when no profiler is listening: a | |
| single atomic load, and the label isn't even built. | |
| KernelAbstractions can put named ranges on the timeline of a tracing profiler, such as | |
| NVIDIA Nsight Systems or Intel VTune, so that you can see which part of your program a | |
| stretch of kernels belongs to. Annotations are cheap when no profiler is listening. |
Comment on lines
+33
to
+37
| Ranges are recorded on the host threads of the process, which is how NVTX, ITT and | ||
| roctx work: it is the profiler that attributes the device work launched within a range to | ||
| it. So which profiler records the ranges depends on what the process runs under, not on the | ||
| backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend | ||
| under VTune gives ITT tasks. |
Member
Author
There was a problem hiding this comment.
Suggested change
| Ranges are recorded on the host threads of the process, which is how NVTX, ITT and | |
| roctx work: it is the profiler that attributes the device work launched within a range to | |
| it. So which profiler records the ranges depends on what the process runs under, not on the | |
| backend: running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend | |
| under VTune gives ITT tasks. | |
| Ranges are recorded on the host threads of the process, which is how NVTX, ITT and | |
| roctx work: it is the profiler that attributes the device work launched within a range to | |
| it. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #836 +/- ##
==========================================
- Coverage 79.06% 0.00% -79.07%
==========================================
Files 24 27 +3
Lines 2040 2075 +35
==========================================
- Hits 1613 0 -1613
- Misses 427 2075 +1648 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Records the ranges, markers and kernel launches of an expression with a temporary tracer, and summarizes them per name or, with `trace = true`, lists them in order. By default kernel launches synchronize their backend while profiling, so that their ranges measure execution rather than launch; tracers opt into this with `synchronizes_launches`. Assisted-by: Claude Code (Opus 5.5)
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
vchuravy
added this pull request to stack #838
October 4, 2026 08:38
Tracers are global, so `@profile` recorded every task in the process, including other profiles. It now records only the task running its expression and the tasks spawned from it, tracked with a ScopedValue, numbers those tasks, and nests the trace per task rather than per thread. It warns about ranges still open when it finishes, i.e. tasks it wasn't waited for. Assisted-by: Claude Code (Opus 5.5)
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a low-cost tracing subsystem that puts named ranges on the timeline of a tracing profiler. This supersedes #703 and #66.
Kernel launches are also recorded as ranges, named after the kernel.
Design
NVTX, ITT and roctx annotate host threads, and the profiler attributes the device work launched within a range to that range. So which profiler records a range depends on what the process runs under, not on the backend. For example, running the CPU backend under Nsight Systems gives NVTX ranges, and a GPU backend under VTune gives ITT tasks. For this reason the API takes no backend argument, and there are no KernelInterface hooks. Every range goes to all registered
KernelAbstractions.Tracers.@profiling_rangeuses a scope-free:tryfinally, like@time. The range is ended if the expression throws, and assignments inside it remain visible afterwards.Tracers
Each tracer registers itself only when its profiler is attached:
NVTXExtNVTX.isactive(), i.e. undernsysROCTXExtrocprofv3(ROCP_TOOL_LIBRARIES) or legacyrocprof(HSA_TOOLS_LIB)IntelITTExtIntelITT.isactive(), i.e. under VTuneNVTXTTracerJULIA_KA_NVTXT=1or=path-%p.nvtxtNVTXTTracerrevives #66. It writes the NVTXT text format, which Nsight Systems can import withImportNvtxt, so you can trace without a profiler attached.Testing
nsys profile --trace=nvtx:nvtx_sumshowsDemo:step 1andKernelAbstractions:scale!ranges./opt/rocm/lib/libroctx64.so(ROCm 7.2) using a stand-inAMDGPUmodule. I could not test it underrocprofv3, which isn't installed.🤖 Generated with Claude Code