Repository navigation
Pytorch backend: analysis - #541
Merged
TApplencourt merged 30 commits intoSep 29, 2026
Merged
Conversation
- libPyTorchInterval.so: babeltrace2 filter turning lttng_ust_pytorch
op_entry/op_exit into generic interval lttng:host messages
- btx_pytorch_model.yaml: hand-written upstream model (cxi-style, no
--matching needed for 2 fixed events)
- btx_pytorchinterval_callbacks.cpp: pairs entry/exit via EntryState,
keyed by {hostname, vpid, vtid}
- Adds BACKEND_PYTORCH to backend_e (utils/xprof_utils.hpp)
Verified against a real trace (854 op_entry/op_exit pairs): to_interval
produces 854 correctly-named, correctly-timed lttng:host messages.
Pending: tally name/level registration, default --backends wiring
(next commit); timeline needs no changes (backend-agnostic, verified).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- utils/xprof_utils.hpp: adds pytorch to pretty_backend_name_g and backend_levels_g at level 6 (its own tier, above itt(5), since RecordFunction ops wrap cuda/ze/omp/mpi calls beneath them) - xprof/xprof.rb.in, utils/babeltrace_thapi.in: add pytorch:6 to the default --backends list so tally/timeline pick it up without an explicit --backends flag Verified against the same real trace (854 op_entry/op_exit pairs): `tally` now prints a dedicated BACKEND_PYTORCH section (854 calls, 27.75ms total) instead of "Wrong Backend passed" warnings; other backends' tally output (ze/cl) is unaffected. Timeline needs no changes -- confirmed backend-agnostic in the prior commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Collaborator
|
Also I think pytorch may work on the |
…pytorch env variable and version validation
Collaborator
|
If you can do one more commit to fix clang-format and rubucopt, and then we merge. : ) Thanks a lot! |
Add a failure-gated step that reproduces the failing iprof command, traces loader binding/relocation with LD_DEBUG, checks libtorch_cpu.so dep resolution, and runs an isolated preload without iprof env-building. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A hung tracee is not a step failure, so the failure-gated diagnostic step never ran. Add timeout-minutes to the integration test to convert a hang into a failure, and wrap the diagnostic iprof calls with timeout so they cannot hang the diagnostic step either. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Detach each probe (setsid + timeout -s KILL) and redirect its stdio to a file so lingering lttng children cannot hold the step pipe open, which previously left the diagnostic step blocked until its timeout with no output. Run under plain bash so a non-zero probe does not abort the step. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GitHub's `shell: bash` keeps `-e -o pipefail`, so the first probe exiting 127 aborted the step before any diagnostic output. Set +e +o pipefail at the top of the script instead. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Removed integration test step and added tmate session setup for debugging.
Removed tmate session setup and related steps from the workflow.
Removed the iprof debug step from the workflow.
Added steps to install PyTorch and update GITHUB_PATH.
Updated PyTorch installation to use the latest version.
TApplencourt
approved these changes
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixed
-Wl,--no-as-needed,-ltorch_cpu,--as-needed.overload_nameto the tracepoints so overloaded ops stay distinct (e.g.aten::empty.memory_formatno longer collapses toaten::empty).pytorch.tp→pytorch_tracepoints.h.import torchfails.LTTNG_UST_PYTORCH_LIBRARY_PATH.Added (analysis)
btx_pytorchinterval_callbacks.cpp) that merges op entry/exit events into completed duration intervals.pytorchbackend (id 10, level 6) with the tally and default--backendslists soiprofproduces the aggregated op table and timeline.btx_pytorch_model.yaml.Added (CI)
batsintegration test that tracestorch.empty(3)and assertsaten::empty.memory_formatappears in the tally.torch==2.14.0install step to the presubmit workflow so the integration test runs on every PR.