Skip to content

[pull] master from tensorflow:master - #253

Open
pull[bot] wants to merge 10000 commits into
greedforgood:masterfrom
tensorflow:master
Open

pull[bot] wants to merge 10000 commits into
greedforgood:masterfrom
tensorflow:master

Conversation

@pull

@pull pull Bot commented Jun 10, 2021 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot]

Can you help keep this open source service alive? 💖 Please sponsor : )

Emilio Cota and others added 27 commits September 29, 2026 16:36
So that all generated code is instrumented.

This paves the way for the upcoming msan support.

PiperOrigin-RevId: 990569171
…ops.

Inside a while loop carrying xla_disable_while_loop_copies, XLA is not free to
insert a relayout copy. This change exposes IsWhileLoopCopyDisabled on
ComputationLayoutConstraints and threads it to the TPU convolution output
layout tie-break to prevent inserting relayout copies inside copy-disabled while loops.

PiperOrigin-RevId: 990570717
…ted` kernel.

- Add `FuseA4W2DRQFullyConnectedPass` to collapse blockwise Q/DQ patterns (symmetric 32-element i4/e8m0 dynamic activations + centered per-channel i2 weights) into a single `tfl.fully_connected` carrying `tfl.quant_spec = {spec = "cint2_fp32_int4_e8m0_drq", act_dilation = ...}` and a per-axis `i2` `tfl.pseudo_qconst`.

- Implement `ParseQuantSpec` and `EvalA4W2DRQ` in the `FullyConnected` reference kernel (bumping max version to 15) to evaluate `a4w2_drq_v1` and reject unrecognized `quant_spec` payloads.

PiperOrigin-RevId: 990578615
Based on https://en.cppreference.com/cpp/types/numeric_limits:

`::min()` returns "the smallest positive normal value of the given
floating-point type" while `::lowest()` returns "the lowest finite value".

PiperOrigin-RevId: 990578654
Ignoring returned future is always an error

Reverts 543710f

PiperOrigin-RevId: 990580556
PiperOrigin-RevId: 990590739
`Subgraph::Prepare` and `Subgraph::Invoke` acquire `Delegate::workspace_mutex_` to serialize access to the shared `xnn_workspace` and its intrusive `first_user` linked list of `xnn_runtime` instances. However, `Subgraph::Create` (`xnn_create_runtime_v4`), `Subgraph::~Subgraph` (`xnn_delete_runtime`), and `Delegate::~Delegate` (`xnn_release_workspace`) previously mutated `workspace->first_user` and `workspace->ref_count` or freed the `xnn_runtime` without holding `workspace_mutex_`.

Share `workspace_mutex_` via `std::shared_ptr<std::mutex>` between `Delegate` and `Subgraph` (matching the ref-counted lifetime of `xnn_workspace`), acquire `workspace_mutex_` around `xnn_create_runtime_v4` in `Subgraph::Create`, before resetting `runtime_` in `Subgraph::~Subgraph`, and before resetting `workspace_` in `Delegate::~Delegate`, clean up `runtime_ptr` under `workspace_mutex_` if `StopBuildStep()` fails, and log an error and return `kTfLiteError` if `runtime_` is null in `Subgraph::Prepare` and `Subgraph::Invoke`.

PiperOrigin-RevId: 990599327
Integrate cl/983398077 (2e332453d92) removed the callers mentioned
in the comment. Remove it.

PiperOrigin-RevId: 990604835
…lation

This change introduces ALG_DOT_BF16_BF16_FP8X3 and ALG_DOT_BF16_BF16_FP8X4 to PrecisionConfig in xla_data.proto and plumbs them through StableHLO and MHLO attribute translators.

PiperOrigin-RevId: 990613954
and restore input schedule when we change schduler config.

PiperOrigin-RevId: 990614610
So that sanitizers won't preclude LLVM passes from combining
them into an fma. Sanitizer instrumentation can break basic blocks,
which prevents LLVM from fusing FMA's since FMA's must come from
same-basic-block pairs.

This change is necessary to maintain the same numerics when we
have full msan support, which is upcoming.

PiperOrigin-RevId: 990627586
… np.uint32

np_test.py imports numpy as 'onp', so np.uint32 raised NameError at test
collection. This matches the fix requested in the review of PR #127683.
…SelectOp` with the index tie-breaking condition rather than `MaxOp`. While `SelectOp` produces the same value for normal numbers, it has different IEEE-754 semantics for `NaN`s (dropping `NaN` when it appears on the LHS) and signed zeros, which caused divergent `NaN` propagation behavior and prevented XLA from simplifying unused-index reductions to `kMaximum`. Rely on `MaxOp` for `selected_value` while keeping `SelectOp` for `selected_index`, and add a unit test in `StablehloBuilderTest` verifying the generated reducer body.

PiperOrigin-RevId: 990642043
… by value so it is copied while the lock is held.

PiperOrigin-RevId: 990657976
Adds the llvm_xz (v5.8.3) external archive in WORKSPACE.bazel and provides third_party/xz.BUILD to build the liblzma library.

PiperOrigin-RevId: 990674437
…in YNNPACK.

Instead of folding query heads into the row/sequence dimension for GQA during prefill, split Q into [n_kv, g] and expand K, V, and the mask to 5D so they broadcast across the group dimension. This updates the K and V transposes to 5D, avoids splitting and fusing logits around mask addition, and fuses the [n_kv, g] head dimensions back together after the P @ V matmul. GQA folding is retained for decode paths and sequence-major inputs.

PiperOrigin-RevId: 990686744
PiperOrigin-RevId: 990706829
Levon Ter-Grigoryan and others added 30 commits October 2, 2026 09:20
In `NonMaxSuppressionMultiClassRegularHelper`,
`box_info_after_regular_non_max_suppression` is sized to
`max_detections + num_detections_per_class`. During the multi-threaded merge
phase, both `sorted_indices_size` and `tasks[j].sorted_indices_size` can reach
`max_detections`. When `num_detections_per_class < max_detections`, appending
all `tasks[j].sorted_indices_size` elements via `memcpy` before calling
`InplaceMergeBoxInfo` overflows the destination buffer on the heap.

Size the buffer to `2 * max_detections`, which covers the multi-threaded merge
and is never smaller than what the single-threaded path needs, because
`num_detections_per_class` is already capped at `max_detections`.

PiperOrigin-RevId: 992360565
…rialization.

The field is optional for now: executables serialized before this change don't carry it, and must remain loadable for the AOT backward-compatibility window.

PiperOrigin-RevId: 992370223
…CompileOptions

`FunctionalHloRunner::Compile` already supports ahead-of-time compilation from a
`PjRtTopologyDescription`, but the `CompileOptions` it needs can only be built
from a `PjRtClient`. This makes it impossible to prepare and compile a program
for a target on a machine that does not have the target devices.

`CreateCompileOptions` only uses the client for two things:

*   `device_count()`, to infer the number of replicas/partitions when neither
    `RawCompileOptions` nor `ExecutionOptions` specify them.
*   `GetDefaultDeviceAssignment(num_replicas, num_partitions)`, when no device
    assignment is provided and `num_slices` is not set.

This change moves the existing body into an internal helper that takes these
two as parameters, and adds a new overload that takes a
`PjRtTopologyDescription` instead of a client:

*   The device count is `topology.DeviceDescriptions().size()`.
*   The default device assignment is
    `topology.GetDefaultDeviceAssignment(task_id, num_replicas, std::nullopt,
    num_partitions, nullptr)`. This matches what the client overload does for
    `CommonPjRtClient`, with `task_id` as the process index.

The existing client overload is behaviorally unchanged.

Tested: new `CreateCompileOptionsFromTopologyMatchesClient` test in
`xla/tools/multihost_hlo_runner/functional_hlo_runner_test.cc`, checking that
both overloads produce the same replicas, partitions and default device
assignment.
PiperOrigin-RevId: 992431526
… allow host compilation after freeze.

PiperOrigin-RevId: 992436199
Updates LLVM usage to match
[018a9e4a74ba](llvm/llvm-project@018a9e4)

PiperOrigin-RevId: 992441813
…Buffer outputs.

PiperOrigin-RevId: 992462398
…Connected kernel.

PiperOrigin-RevId: 992533092
…ple_args is enabled.

PiperOrigin-RevId: 992547917
…ction and update hlo_dump_ui_bin_sanitized.js golden file.

PiperOrigin-RevId: 992549561
… fusions.

When a library fusion (e.g. YNNPACK) terminates with a data movement,
reshaping, or broadcasting operation (broadcast, reshape, bitcast, transpose,
copy, or pad), peel that instruction out of the fusion into the parent
computation.

Terminating a library fusion with a broadcast forces libraries to materialize
the broadcast into memory, whereas keeping it outside allows downstream
consumers to fuse it on-the-fly.

Impact to recent benchmarks from openxla/xla#49131:
```
name                                                                                   cpu/op         cpu/op       vs base
BM_CompileHloModule/pairwise_biot_savart_4096_f64/process_time                          94.13m ±  2%    93.41m ±   1%        ~ (p=0.137 n=15)
BM_CompileHloModule/pairwise_biot_savart_grad_2048_f64/process_time                     336.9m ±  1%    349.8m ±   1%   +3.85% (p=0.000 n=15)
BM_CompileHloModule/pairwise_gravity_4096_f64/process_time                              5.189m ±  3%    5.241m ±   2%        ~ (p=0.461 n=15)
BM_CompileHloModule/pairwise_gravity_grad_2048_f64/process_time                         8.452m ±  4%    8.407m ±   4%        ~ (p=0.870 n=15)
BM_CompileHloModule/pairwise_min_distance_grad_2048_f64/process_time                    146.6m ±  1%    116.2m ±   1%  -20.74% (p=0.000 n=15)
BM_CompileHloModule/pairwise_rbf_gradient_4096_f32/process_time                         4.958m ±  3%    4.988m ±   4%        ~ (p=0.935 n=15)
BM_CompileHloModule/pairwise_rbf_gradient_grad_2048_f32/process_time                    103.5m ±  1%    124.7m ±   2%  +20.45% (p=0.000 n=15)
BM_CompileHloModule/pairwise_sum_distances_grad_2048_f32/process_time                   78.02m ±  2%    78.53m ±   1%        ~ (p=0.838 n=15)

name                                                                                   time/op        time/op     vs base
BM_HloModule/pairwise_biot_savart_4096_f64/process_time                                101.90m ±  3%    52.42m ±  4%  -48.56% (p=0.000 n=15)
BM_HloModule/pairwise_biot_savart_grad_2048_f64/process_time                            73.91m ±  2%    60.11m ±  5%  -18.68% (p=0.000 n=15)
BM_HloModule/pairwise_gravity_4096_f64/process_time                                     41.10m ±  1%    41.31m ±  1%        ~ (p=0.098 n=15)
BM_HloModule/pairwise_gravity_grad_2048_f64/process_time                                25.69m ±  1%    25.88m ±  1%        ~ (p=0.106 n=15)
BM_HloModule/pairwise_min_distance_grad_2048_f64/process_time                           17.00m ±  4%    13.55m ±  1%  -20.29% (p=0.000 n=15)
BM_HloModule/pairwise_rbf_gradient_4096_f32/process_time                                20.30m ±  1%    20.26m ±  2%        ~ (p=0.436 n=15)
BM_HloModule/pairwise_rbf_gradient_grad_2048_f32/process_time                           34.14m ±  2%    18.69m ±  2%  -45.24% (p=0.000 n=15)
BM_HloModule/pairwise_sum_distances_grad_2048_f32/process_time                          7.021m ±  4%    7.065m ±  3%        ~ (p=0.486 n=15)
```

PiperOrigin-RevId: 992593366
… date attribute.

PiperOrigin-RevId: 992601667
…ed_4bit FullyConnected kernel.

Reverts 661ea39

PiperOrigin-RevId: 992618514
CommonPjRtClient::ShouldDoDirectTransfer and remove PjRtStreamExecutorClient.

PiperOrigin-RevId: 992629795
… to mhlo.spmd_parameters_shardings and mhlo.spmd_output_sharding.

Add tests.

PiperOrigin-RevId: 992646101
PiperOrigin-RevId: 992706353
PiperOrigin-RevId: 992711490
With the compute-synchronized allocation model, an allocated buffer captures
the next compute stream sync point, but the corresponding event is recorded
lazily at the tail of the compute stream when `WaitForAllocation` runs. Every
transfer path defers `WaitForAllocation` to `async_work_runner()` (until
transfer dependencies, remote descriptors, etc. are ready), so an execution
enqueued on the compute stream in the meantime ends up ahead of the allocation
event. The transfer then falsely depends on that execution, which can deadlock
when the execution is a collective whose peer is waiting on the transfer.

This change adds `PjRtStreamExecutorRawClient::MaterializeAllocationEvent()`
and calls it on the scheduling thread before deferring in
`ScheduleRemoteSend`, `CrossHostReceiveBuffersInto`,
`CopyRawHostToDeviceAndReturnEvent`, `CopyRawDeviceToHostAndReturnEvent`,
`ScheduleCopyTo` and `IntraClientCopyToWithDependencies`. The deferred
`WaitForAllocation` then reuses the already recorded event.

PiperOrigin-RevId: 992717004
In Confidential Computing VMs (e.g. CTIVM + GPU CC), userspace host memory is private by default and inaccessible by the GPU, which causes `cuMemHostRegister` to fail. PjRt GPU must stage transfers onto host memory allocated using `cuMemHostAlloc`.

1. Query NVML for Confidential Computing status (`nvmlSystemGetConfComputeState`) in `stream_executor/cuda` and plumb it via `DeviceDescription`, `GpuDeviceInfoProto`, and `GpuTopology` through `StreamExecutorGpuTopologyDescription`.
2. When Confidential Computing mode is enabled:
   - Make `DmaMap` and `DmaUnmap` no-ops in `StreamExecutorGpuRawClient`.
   - Override `ShouldStageHostToDeviceTransfers` in `StreamExecutorGpuRawClient` to always return true.
   - Fail PjRt GPU client creation if a custom host memory allocator is enabled.
   - In `CudaHostAllocator`, bypass NUMA allocation and directly allocate via `cuMemHostAlloc` and free via `cuMemFreeHost`.

PiperOrigin-RevId: 992778826
PiperOrigin-RevId: 992845247
HloReachabilityMap::Build zero filled the n by n bit matrix and then
overwrote every row in post order. For a 250k instruction computation
that is a full pass over 7.8 GB before the build starts. Build now
allocates the rows without zero filling them and writes each row once:
the first input row is copied in, the rest are unioned, input free rows
are cleared, and the diagonal bit is set last. SetReachabilityToUnion
copies the first input row the same way. The bits are unchanged.
PiperOrigin-RevId: 992957894
… the ptr inside the range.

In this way, `exported_fabric_handles_` (keyed by {executor, ptr}) will cache all exports within the underlying allocation. And on the import side, the process will be skipped because of the imported_fabric_handles_ cache.

PiperOrigin-RevId: 993071288
cl/992717004 started calling `MaterializeAllocationEvent(*raw_buffer)` at the
top of `StreamExecutorGpuRawClient::ScheduleRemoteSend`, but `raw_buffer` is
null when the source buffer is an error buffer (the error is reported through
`definition_events` in the deferred callback). This crashed
`CrossHostTransferSourceBufferError`. Skip materialization for null buffers.

PiperOrigin-RevId: 993113862
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.