Skip to content

Pytorch backend: analysis - #541

Merged
TApplencourt merged 30 commits into
argonne-lcf:develfrom
DonAurelio:pytorch-backend-analysis
Sep 29, 2026
Merged

TApplencourt merged 30 commits into
argonne-lcf:develfrom
DonAurelio:pytorch-backend-analysis

Conversation

@DonAurelio

@DonAurelio DonAurelio commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Fixed

  • Fixed entry/exit duration corruption on nested/reentrant ATen ops.
  • Fixed the tracer failing to load in CI (undefined symbol: at::addGlobalCallback) on --as-needed toolchains, via -Wl,--no-as-needed,-ltorch_cpu,--as-needed.
  • Added overload_name to the tracepoints so overloaded ops stay distinct (e.g. aten::empty.memory_format no longer collapses to aten::empty).
  • Renamed the tracepoint provider pytorch.tp → pytorch_tracepoints.h.
  • Changed torch-library discovery to warn instead of aborting when import torch fails.
  • Added a tested-version check (1.11.0–2.14.0) that warns on untested PyTorch releases.
  • Added LTTNG_UST_PYTORCH_LIBRARY_PATH.

Added (analysis)

  • Added the interval filter (btx_pytorchinterval_callbacks.cpp) that merges op entry/exit events into completed duration intervals.
  • Registered the pytorch backend (id 10, level 6) with the tally and default --backends lists so iprof produces the aggregated op table and timeline.
  • Added the btx_pytorch_model.yaml.

Added (CI)

  • Added a bats integration test that traces torch.empty(3) and asserts aten::empty.memory_format appears in the tally.
  • Added a CPU-only torch==2.14.0 install step to the presubmit workflow so the integration test runs on every PR.

Aurelio Vivas and others added 6 commits September 14, 2026 23:48
- libPyTorchInterval.so: babeltrace2 filter turning lttng_ust_pytorch
  op_entry/op_exit into generic interval lttng:host messages
- btx_pytorch_model.yaml: hand-written upstream model (cxi-style, no
  --matching needed for 2 fixed events)
- btx_pytorchinterval_callbacks.cpp: pairs entry/exit via EntryState,
  keyed by {hostname, vpid, vtid}
- Adds BACKEND_PYTORCH to backend_e (utils/xprof_utils.hpp)

Verified against a real trace (854 op_entry/op_exit pairs): to_interval
produces 854 correctly-named, correctly-timed lttng:host messages.

Pending: tally name/level registration, default --backends wiring
(next commit); timeline needs no changes (backend-agnostic, verified).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- utils/xprof_utils.hpp: adds pytorch to pretty_backend_name_g and
  backend_levels_g at level 6 (its own tier, above itt(5), since
  RecordFunction ops wrap cuda/ze/omp/mpi calls beneath them)
- xprof/xprof.rb.in, utils/babeltrace_thapi.in: add pytorch:6 to the
  default --backends list so tally/timeline pick it up without an
  explicit --backends flag

Verified against the same real trace (854 op_entry/op_exit pairs):
`tally` now prints a dedicated BACKEND_PYTORCH section (854 calls,
27.75ms total) instead of "Wrong Backend passed" warnings; other
backends' tally output (ze/cl) is unaffected.

Timeline needs no changes -- confirmed backend-agnostic in the prior
commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Comment thread backends/pytorch/include/ATen/record_function.h Outdated
Comment thread backends/pytorch/btx_pytorchinterval_callbacks.cpp Outdated
Comment thread backends/pytorch/tracer_pytorch.cpp Outdated
Comment thread backends/pytorch/tracer_pytorch.cpp Outdated
Comment thread xprof/xprof.rb.in Outdated
Comment thread xprof/xprof.rb.in Outdated
Comment thread xprof/xprof.rb.in Outdated
Comment thread xprof/xprof.rb.in Outdated
Comment thread xprof/xprof.rb.in Outdated
Comment thread xprof/xprof.rb.in Outdated
Comment thread xprof/xprof.rb.in Outdated
@TApplencourt

Copy link
Copy Markdown
Collaborator

Also I think pytorch may work on the github-CI machine? So we can add a integration tests if you don't mind. :)

Comment thread backends/pytorch/tracer_pytorch.cpp Outdated
Comment thread xprof/xprof.rb.in Outdated
@TApplencourt

Copy link
Copy Markdown
Collaborator

If you can do one more commit to fix clang-format and rubucopt, and then we merge. : )


1 file inspected, no offenses detected
Inspecting 1 file
C

Offenses:

xprof/xprof.rb.in:156:15: C: [Correctable] Style/TrailingUnderscoreVariable: Do not use trailing _s in parallel assignment. Prefer stdout_str, = exec("#{whichlib64_bin} #{binary} #{libs.join(' ')}",
                    opts: { 'LD_LIBRARY_PATH' => ld_path },
                    ignore_exit_codes: [1, 2]).
  stdout_str, _, _ = exec("#{whichlib64_bin} #{binary} #{libs.join(' ')}",
              ^^^^^
xprof/xprof.rb.in:157:21: C: [Correctable] Layout/ArgumentAlignment: Align the arguments of a method call if they span more than one line.
                    opts: { 'LD_LIBRARY_PATH' => ld_path }, ...
                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
xprof/xprof.rb.in:895:26: C: [Correctable] Layout/ExtraSpacing: Unnecessary spacing detected.
    stdout_str, _, status  = exec(torch_probe, ignore_exit_codes: [1, 127])
                         ^

1 file inspected, 3 offenses detected, 3 offenses autocorrectable

Thanks a lot!

DonAurelio and others added 14 commits September 29, 2026 14:00
Add a failure-gated step that reproduces the failing iprof command,
traces loader binding/relocation with LD_DEBUG, checks libtorch_cpu.so
dep resolution, and runs an isolated preload without iprof env-building.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A hung tracee is not a step failure, so the failure-gated diagnostic
step never ran. Add timeout-minutes to the integration test to convert a
hang into a failure, and wrap the diagnostic iprof calls with timeout so
they cannot hang the diagnostic step either.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Detach each probe (setsid + timeout -s KILL) and redirect its stdio to a
file so lingering lttng children cannot hold the step pipe open, which
previously left the diagnostic step blocked until its timeout with no
output. Run under plain bash so a non-zero probe does not abort the step.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GitHub's `shell: bash` keeps `-e -o pipefail`, so the first probe exiting
127 aborted the step before any diagnostic output. Set +e +o pipefail at
the top of the script instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Removed integration test step and added tmate session setup for debugging.
Removed tmate session setup and related steps from the workflow.
Removed the iprof debug step from the workflow.
Added steps to install PyTorch and update GITHUB_PATH.
Comment thread .github/workflows/presubmit.yml Outdated
Comment thread backends/pytorch/Makefile.am
Updated PyTorch installation to use the latest version.

@nscottnichols nscottnichols left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@TApplencourt
TApplencourt merged commit 6d0fd2c into argonne-lcf:devel Sep 29, 2026
14 checks passed
@DonAurelio
DonAurelio deleted the pytorch-backend-analysis branch September 30, 2026 11:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants