Conversation
So that all generated code is instrumented. This paves the way for the upcoming msan support. PiperOrigin-RevId: 990569171
…ops. Inside a while loop carrying xla_disable_while_loop_copies, XLA is not free to insert a relayout copy. This change exposes IsWhileLoopCopyDisabled on ComputationLayoutConstraints and threads it to the TPU convolution output layout tie-break to prevent inserting relayout copies inside copy-disabled while loops. PiperOrigin-RevId: 990570717
…Backend PiperOrigin-RevId: 990574854
…ted` kernel.
- Add `FuseA4W2DRQFullyConnectedPass` to collapse blockwise Q/DQ patterns (symmetric 32-element i4/e8m0 dynamic activations + centered per-channel i2 weights) into a single `tfl.fully_connected` carrying `tfl.quant_spec = {spec = "cint2_fp32_int4_e8m0_drq", act_dilation = ...}` and a per-axis `i2` `tfl.pseudo_qconst`.
- Implement `ParseQuantSpec` and `EvalA4W2DRQ` in the `FullyConnected` reference kernel (bumping max version to 15) to evaluate `a4w2_drq_v1` and reject unrecognized `quant_spec` payloads.
PiperOrigin-RevId: 990578615
Based on https://en.cppreference.com/cpp/types/numeric_limits: `::min()` returns "the smallest positive normal value of the given floating-point type" while `::lowest()` returns "the lowest finite value". PiperOrigin-RevId: 990578654
Ignoring returned future is always an error Reverts 543710f PiperOrigin-RevId: 990580556
PiperOrigin-RevId: 990583743
PiperOrigin-RevId: 990590739
`Subgraph::Prepare` and `Subgraph::Invoke` acquire `Delegate::workspace_mutex_` to serialize access to the shared `xnn_workspace` and its intrusive `first_user` linked list of `xnn_runtime` instances. However, `Subgraph::Create` (`xnn_create_runtime_v4`), `Subgraph::~Subgraph` (`xnn_delete_runtime`), and `Delegate::~Delegate` (`xnn_release_workspace`) previously mutated `workspace->first_user` and `workspace->ref_count` or freed the `xnn_runtime` without holding `workspace_mutex_`. Share `workspace_mutex_` via `std::shared_ptr<std::mutex>` between `Delegate` and `Subgraph` (matching the ref-counted lifetime of `xnn_workspace`), acquire `workspace_mutex_` around `xnn_create_runtime_v4` in `Subgraph::Create`, before resetting `runtime_` in `Subgraph::~Subgraph`, and before resetting `workspace_` in `Delegate::~Delegate`, clean up `runtime_ptr` under `workspace_mutex_` if `StopBuildStep()` fails, and log an error and return `kTfLiteError` if `runtime_` is null in `Subgraph::Prepare` and `Subgraph::Invoke`. PiperOrigin-RevId: 990599327
Integrate cl/983398077 (2e332453d92) removed the callers mentioned in the comment. Remove it. PiperOrigin-RevId: 990604835
…ror builders PiperOrigin-RevId: 990607199
…lation This change introduces ALG_DOT_BF16_BF16_FP8X3 and ALG_DOT_BF16_BF16_FP8X4 to PrecisionConfig in xla_data.proto and plumbs them through StableHLO and MHLO attribute translators. PiperOrigin-RevId: 990613954
and restore input schedule when we change schduler config. PiperOrigin-RevId: 990614610
So that sanitizers won't preclude LLVM passes from combining them into an fma. Sanitizer instrumentation can break basic blocks, which prevents LLVM from fusing FMA's since FMA's must come from same-basic-block pairs. This change is necessary to maintain the same numerics when we have full msan support, which is upcoming. PiperOrigin-RevId: 990627586
… np.uint32 np_test.py imports numpy as 'onp', so np.uint32 raised NameError at test collection. This matches the fix requested in the review of PR #127683.
…SelectOp` with the index tie-breaking condition rather than `MaxOp`. While `SelectOp` produces the same value for normal numbers, it has different IEEE-754 semantics for `NaN`s (dropping `NaN` when it appears on the LHS) and signed zeros, which caused divergent `NaN` propagation behavior and prevented XLA from simplifying unused-index reductions to `kMaximum`. Rely on `MaxOp` for `selected_value` while keeping `SelectOp` for `selected_index`, and add a unit test in `StablehloBuilderTest` verifying the generated reducer body. PiperOrigin-RevId: 990642043
… by value so it is copied while the lock is held. PiperOrigin-RevId: 990657976
PiperOrigin-RevId: 990670493
Adds the llvm_xz (v5.8.3) external archive in WORKSPACE.bazel and provides third_party/xz.BUILD to build the liblzma library. PiperOrigin-RevId: 990674437
…in YNNPACK. Instead of folding query heads into the row/sequence dimension for GQA during prefill, split Q into [n_kv, g] and expand K, V, and the mask to 5D so they broadcast across the group dimension. This updates the K and V transposes to 5D, avoids splitting and fusing logits around mask addition, and fuses the [n_kv, g] head dimensions back together after the P @ V matmul. GQA folding is retained for decode paths and sequence-major inputs. PiperOrigin-RevId: 990686744
PiperOrigin-RevId: 990706829
PiperOrigin-RevId: 992354817
In `NonMaxSuppressionMultiClassRegularHelper`, `box_info_after_regular_non_max_suppression` is sized to `max_detections + num_detections_per_class`. During the multi-threaded merge phase, both `sorted_indices_size` and `tasks[j].sorted_indices_size` can reach `max_detections`. When `num_detections_per_class < max_detections`, appending all `tasks[j].sorted_indices_size` elements via `memcpy` before calling `InplaceMergeBoxInfo` overflows the destination buffer on the heap. Size the buffer to `2 * max_detections`, which covers the multi-threaded merge and is never smaller than what the single-threaded path needs, because `num_detections_per_class` is already capped at `max_detections`. PiperOrigin-RevId: 992360565
…rialization. The field is optional for now: executables serialized before this change don't carry it, and must remain loadable for the AOT backward-compatibility window. PiperOrigin-RevId: 992370223
…uilds. PiperOrigin-RevId: 992395219
…opology. PiperOrigin-RevId: 992420737
…CompileOptions
`FunctionalHloRunner::Compile` already supports ahead-of-time compilation from a
`PjRtTopologyDescription`, but the `CompileOptions` it needs can only be built
from a `PjRtClient`. This makes it impossible to prepare and compile a program
for a target on a machine that does not have the target devices.
`CreateCompileOptions` only uses the client for two things:
* `device_count()`, to infer the number of replicas/partitions when neither
`RawCompileOptions` nor `ExecutionOptions` specify them.
* `GetDefaultDeviceAssignment(num_replicas, num_partitions)`, when no device
assignment is provided and `num_slices` is not set.
This change moves the existing body into an internal helper that takes these
two as parameters, and adds a new overload that takes a
`PjRtTopologyDescription` instead of a client:
* The device count is `topology.DeviceDescriptions().size()`.
* The default device assignment is
`topology.GetDefaultDeviceAssignment(task_id, num_replicas, std::nullopt,
num_partitions, nullptr)`. This matches what the client overload does for
`CommonPjRtClient`, with `task_id` as the process index.
The existing client overload is behaviorally unchanged.
Tested: new `CreateCompileOptionsFromTopologyMatchesClient` test in
`xla/tools/multihost_hlo_runner/functional_hlo_runner_test.cc`, checking that
both overloads produce the same replicas, partitions and default device
assignment.
PiperOrigin-RevId: 992431526
… allow host compilation after freeze. PiperOrigin-RevId: 992436199
Updates LLVM usage to match [018a9e4a74ba](llvm/llvm-project@018a9e4) PiperOrigin-RevId: 992441813
…Buffer outputs. PiperOrigin-RevId: 992462398
…Connected kernel. PiperOrigin-RevId: 992533092
…ple_args is enabled. PiperOrigin-RevId: 992547917
…ction and update hlo_dump_ui_bin_sanitized.js golden file. PiperOrigin-RevId: 992549561
PiperOrigin-RevId: 992582023
… fusions. When a library fusion (e.g. YNNPACK) terminates with a data movement, reshaping, or broadcasting operation (broadcast, reshape, bitcast, transpose, copy, or pad), peel that instruction out of the fusion into the parent computation. Terminating a library fusion with a broadcast forces libraries to materialize the broadcast into memory, whereas keeping it outside allows downstream consumers to fuse it on-the-fly. Impact to recent benchmarks from openxla/xla#49131: ``` name cpu/op cpu/op vs base BM_CompileHloModule/pairwise_biot_savart_4096_f64/process_time 94.13m ± 2% 93.41m ± 1% ~ (p=0.137 n=15) BM_CompileHloModule/pairwise_biot_savart_grad_2048_f64/process_time 336.9m ± 1% 349.8m ± 1% +3.85% (p=0.000 n=15) BM_CompileHloModule/pairwise_gravity_4096_f64/process_time 5.189m ± 3% 5.241m ± 2% ~ (p=0.461 n=15) BM_CompileHloModule/pairwise_gravity_grad_2048_f64/process_time 8.452m ± 4% 8.407m ± 4% ~ (p=0.870 n=15) BM_CompileHloModule/pairwise_min_distance_grad_2048_f64/process_time 146.6m ± 1% 116.2m ± 1% -20.74% (p=0.000 n=15) BM_CompileHloModule/pairwise_rbf_gradient_4096_f32/process_time 4.958m ± 3% 4.988m ± 4% ~ (p=0.935 n=15) BM_CompileHloModule/pairwise_rbf_gradient_grad_2048_f32/process_time 103.5m ± 1% 124.7m ± 2% +20.45% (p=0.000 n=15) BM_CompileHloModule/pairwise_sum_distances_grad_2048_f32/process_time 78.02m ± 2% 78.53m ± 1% ~ (p=0.838 n=15) name time/op time/op vs base BM_HloModule/pairwise_biot_savart_4096_f64/process_time 101.90m ± 3% 52.42m ± 4% -48.56% (p=0.000 n=15) BM_HloModule/pairwise_biot_savart_grad_2048_f64/process_time 73.91m ± 2% 60.11m ± 5% -18.68% (p=0.000 n=15) BM_HloModule/pairwise_gravity_4096_f64/process_time 41.10m ± 1% 41.31m ± 1% ~ (p=0.098 n=15) BM_HloModule/pairwise_gravity_grad_2048_f64/process_time 25.69m ± 1% 25.88m ± 1% ~ (p=0.106 n=15) BM_HloModule/pairwise_min_distance_grad_2048_f64/process_time 17.00m ± 4% 13.55m ± 1% -20.29% (p=0.000 n=15) BM_HloModule/pairwise_rbf_gradient_4096_f32/process_time 20.30m ± 1% 20.26m ± 2% ~ (p=0.436 n=15) BM_HloModule/pairwise_rbf_gradient_grad_2048_f32/process_time 34.14m ± 2% 18.69m ± 2% -45.24% (p=0.000 n=15) BM_HloModule/pairwise_sum_distances_grad_2048_f32/process_time 7.021m ± 4% 7.065m ± 3% ~ (p=0.486 n=15) ``` PiperOrigin-RevId: 992593366
… date attribute. PiperOrigin-RevId: 992601667
…ace. PiperOrigin-RevId: 992602648
…ed_4bit FullyConnected kernel. Reverts 661ea39 PiperOrigin-RevId: 992618514
CommonPjRtClient::ShouldDoDirectTransfer and remove PjRtStreamExecutorClient. PiperOrigin-RevId: 992629795
PiperOrigin-RevId: 992641020
… to mhlo.spmd_parameters_shardings and mhlo.spmd_output_sharding. Add tests. PiperOrigin-RevId: 992646101
[List of integrated commits](triton-lang/triton@a77e7c7...2074a1b) PiperOrigin-RevId: 992675096
PiperOrigin-RevId: 992706353
PiperOrigin-RevId: 992711490
With the compute-synchronized allocation model, an allocated buffer captures the next compute stream sync point, but the corresponding event is recorded lazily at the tail of the compute stream when `WaitForAllocation` runs. Every transfer path defers `WaitForAllocation` to `async_work_runner()` (until transfer dependencies, remote descriptors, etc. are ready), so an execution enqueued on the compute stream in the meantime ends up ahead of the allocation event. The transfer then falsely depends on that execution, which can deadlock when the execution is a collective whose peer is waiting on the transfer. This change adds `PjRtStreamExecutorRawClient::MaterializeAllocationEvent()` and calls it on the scheduling thread before deferring in `ScheduleRemoteSend`, `CrossHostReceiveBuffersInto`, `CopyRawHostToDeviceAndReturnEvent`, `CopyRawDeviceToHostAndReturnEvent`, `ScheduleCopyTo` and `IntraClientCopyToWithDependencies`. The deferred `WaitForAllocation` then reuses the already recorded event. PiperOrigin-RevId: 992717004
PiperOrigin-RevId: 992746570
In Confidential Computing VMs (e.g. CTIVM + GPU CC), userspace host memory is private by default and inaccessible by the GPU, which causes `cuMemHostRegister` to fail. PjRt GPU must stage transfers onto host memory allocated using `cuMemHostAlloc`. 1. Query NVML for Confidential Computing status (`nvmlSystemGetConfComputeState`) in `stream_executor/cuda` and plumb it via `DeviceDescription`, `GpuDeviceInfoProto`, and `GpuTopology` through `StreamExecutorGpuTopologyDescription`. 2. When Confidential Computing mode is enabled: - Make `DmaMap` and `DmaUnmap` no-ops in `StreamExecutorGpuRawClient`. - Override `ShouldStageHostToDeviceTransfers` in `StreamExecutorGpuRawClient` to always return true. - Fail PjRt GPU client creation if a custom host memory allocator is enabled. - In `CudaHostAllocator`, bypass NUMA allocation and directly allocate via `cuMemHostAlloc` and free via `cuMemFreeHost`. PiperOrigin-RevId: 992778826
PiperOrigin-RevId: 992845247
HloReachabilityMap::Build zero filled the n by n bit matrix and then overwrote every row in post order. For a 250k instruction computation that is a full pass over 7.8 GB before the build starts. Build now allocates the rows without zero filling them and writes each row once: the first input row is copied in, the rest are unioned, input free rows are cleared, and the diagonal bit is set last. SetReachabilityToUnion copies the first input row the same way. The bits are unchanged. PiperOrigin-RevId: 992957894
… the ptr inside the range.
In this way, `exported_fabric_handles_` (keyed by {executor, ptr}) will cache all exports within the underlying allocation. And on the import side, the process will be skipped because of the imported_fabric_handles_ cache.
PiperOrigin-RevId: 993071288
cl/992717004 started calling `MaterializeAllocationEvent(*raw_buffer)` at the top of `StreamExecutorGpuRawClient::ScheduleRemoteSend`, but `raw_buffer` is null when the source buffer is an error buffer (the error is reported through `definition_events` in the deferred callback). This crashed `CrossHostTransferSourceBufferError`. Skip materialization for null buffers. PiperOrigin-RevId: 993113862
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot]
Can you help keep this open source service alive? 💖 Please sponsor : )