Conversation
vchuravy
added this pull request to stack #838
October 4, 2026 08:38
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## vc/tracing-v2 #837 +/- ##
=================================================
+ Coverage 76.32% 80.68% +4.36%
=================================================
Files 29 29
Lines 2547 2496 -51
=================================================
+ Hits 1944 2014 +70
+ Misses 603 482 -121 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
vchuravy
force-pushed
the
vc/device-timestamps
branch
5 times, most recently
from
October 4, 2026 15:50
3e1a30a to
990d507
Compare
Adds optional KernelInterface hooks, `record_timestamp(backend)` and `elapsed_time(backend, start, stop)`, with a fallback that reports no support. Tracers that set `records_kernels` get a `KernelTimer` for every launch, with timestamps recorded around the native enqueue alone, so compilation isn't counted. On backends without timestamps the launch synchronizes and is timed on the host instead. `@profile` uses this by default (`device = true`), resolving the timestamps once the profiled expression has run, and reports host-side and device-side activity separately. Launches no longer synchronize by default. Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/device-timestamps
branch
from
October 4, 2026 17:39
990d507 to
cf18848
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #836. This is a prototype of what CUDA's
@profilegets from CUPTI: timing kernels on the device without forcing synchronization.Design
KI.record_timestamp(backend)puts a timestamp on the calling task's queue without blocking. It returnsnothingif the backend doesn't support timestamps.KI.elapsed_time(backend, start, stop)::Int64returns the device nanoseconds between two timestamps, waiting cooperatively if needed.records_kernels,launch_kernelrecords timestamps around the native launch only, so compilation isn't counted as device time. The tracer gets aKernelTimerholding the pair.*in the report.@profile(nowdevice = trueby default). It keeps the timestamp pairs and turns them into durations once the profiled expression has finished, the way CUPTI buffers activity records. The report has separate host-side and device-side sections.trace = truemode, kernels are placed on the host timeline relative to the first kernel on each device, assuming that one started when it was launched. This is approximate.CUDA implementation
Not part of this PR. It is local branch
vc/ka-timestampsof CUDA.jl, on top ofka-0.10:Results (Quadro RTX 4000, 4M-element
mul2/add)Device times match CUDA.
CUDA.@elapsedgives a median of 138 µs formul2, and the profiler measures 138 µs. Timing on the host withsynchronize = trueinstead gives 143 µs, which includes sync overhead.Cost per launch, small kernel:
@profile device = false@profileThe extra ~6 µs is most likely from creating two new
CuEvents, with finalizers, on every launch; an event pool would avoid it. Nothing changes outside@profile.Tasks
Each kernel is attributed to the task that launched it, and in
trace = truemode it appears on a per-queue row such asCUDABackend 1, task 2. Kernels running concurrently on different tasks' streams therefore appear side by side, rather than stacked as if nested.As in #836,
@profileonly records the tasks of its expression, and an unrelated task doesn't pay for device timestamps. CUDA can measure elapsed time between events on different streams of the same context, so placing kernels from different streams on the timeline works.Verified on the RTX 4000 with three
KA.@spawntasks, each running its kernels on its own stream:Caveats
main. CUDA.jl'ska-0.10branch still requires LLVM.jl 9, whilemainrequires LLVM.jl 10. The GPU numbers above come from these commits replayed onto2e027f09, the last commit before the LLVM.jl 10 bump. The test suite here passes onmain, where the CPU backend takes the host-timed path.🤖 Generated with Claude Code