refactor: ADR-019 — remove GPU knowledge from core tensor (#126) - #149
Merged
Merged
Conversation
…terface (ADR-019 Phase 1) Add opaque backendData any field to RawTensor with BackendData/SetBackendData accessors. Add Materialize method to Backend interface — CPU returns buffer directly, WebGPU triggers GPU readback, Autodiff delegates to inner backend. Purely additive — no existing fields or behavior removed. Phase 2 will migrate Data() to use Materialize; Phase 4 will remove legacy GPU fields.
Data() now checks for a Materializer before falling back to the legacy gpuData.Realize() path. WebGPU createLazyResult sets both SetBackendData and SetMaterializer so the new path triggers first. After materialization, backendData and materializer are cleared to prevent double-realization. Legacy GPU path is preserved for Phase 3-4 callers.
Remove 6 flaky pre-readback assertions (activeBatchCount, IsRealized) that tested internal GPU state and raced under load. Replace with numerical correctness checks — tests now verify computed values, not implementation details. Post-readback drain checks (count == 0) kept as valid contract assertions.
…9 Phase 3) Add BackendReleaser interface with ReleaseBackendData, SetPersistent, IsPersistent. Implemented on all backends (CPU=no-op, WebGPU=delegates to LazyGPUData, Autodiff=delegates to inner). Migrate 35 callers in autodiff tape/backward, optim adam/sgd, nn.Parameter, and models/llama to dual-write: new BackendReleaser path + legacy ReleaseGPU/SetGPUPersistent preserved for Phase 4.
Delete lazy_gpu.go and lazy_gpu_wasm.go from internal/tensor/. Move LazyGPUData to internal/backend/webgpu/lazy_gpu_data.go. Remove gpuData/gpuPersistDefer fields and GPUData/SetGPUData/ReleaseGPU/ SetGPUPersistent/IsLazy methods from RawTensor. Remove all legacy dual-write calls from autodiff, optim, nn, and models. Data() now uses only the Materializer path for GPU readback. internal/tensor/ has zero build tags, zero GPU knowledge, zero WASM stubs. 20 files changed, -534 lines net reduction.
Delete 5 sigmoid SIMD files from internal/tensor/ (vendored AVX2 kernel predating the cpu backend's SIMD infrastructure). Tensor Sigmoid/SiLU now use scalar loops — the cpu backend's goexperiment.simd kernel handles all training workloads. internal/tensor/ now has zero //go:build tags, zero platform-specific files, zero GPU knowledge. ADR-019 is complete.
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
This was referenced Aug 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements ADR-019 across 5 phases — removes all GPU-specific code from
internal/tensor/, making it fully backend-agnostic.Phase 1:
backendData any+MaterializebackendData anyfield on RawTensorMaterialize(t *RawTensor) ([]byte, error)on Backend interfacePhase 2:
Materializer+ newData()pathMaterializerinterface for lazy GPU readbackData()uses Materializer when set, legacy path as fallbackPhase 3:
BackendReleaser+ caller migrationBackendReleaserinterface (ReleaseBackendData, SetPersistent, IsPersistent)Phase 4: Delete legacy GPU code
lazy_gpu.goandlazy_gpu_wasm.godeleted frominternal/tensor/LazyGPUDatamoved tointernal/backend/webgpu/Phase 5: Remove SIMD from tensor
internal/tensor/now has zero//go:buildtagsAlso included
Results
internal/tensor/: zero build tags, zero platform-specific files, zero GPU knowledgeTest plan
go build ./...GOOS=js GOARCH=wasm go build ./...go test -short ./...— 22 packages, 0 failuresgo vet ./...golangci-lint run— 0 issues