fix: critical GPU training bugs from Fable 5.1 audit (ADR-020) - #156
Merged
Merged
Conversation
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
- AutodiffBackend forwards ClearInputBufferCache to inner backend, fixing silent stale-weight training when optimizer runs through autodiff wrapper (TASK-154) - ReclaimMemory clears input buffer cache as defense-in-depth - Public autodiff.ReleaseGradients accepts optional backend parameter, fixing no-op that never released GPU gradient buffers (TASK-155) - Cache buildOneHotIdentity per (numClasses, dtype) via sync.Map, preventing 4 GB/step allocation for large vocab backward (TASK-156) - Update gogpu deps: wgpu v0.34.4, gputypes v0.8.0, naga v0.19.0 - Fix wgpu API: AdapterInfo/Backends from gputypes, software.NewBackend
tape.Backward now accepts outputTensor parameter to identify the loss operation on the tape. Backward walks from that op, ignoring any ops recorded after loss (metrics, logging). Falls back to lastOp when outputTensor is nil (backward-compatible).
… (ADR-020 F9) Conv2DInputBackward, Conv2DKernelBackward, and MaxPool2DBackward no longer panic on WebGPU backend. They materialize GPU tensors to CPU, delegate to cpu.New(), and return the result. CNN training on GPU now works (with GPU→CPU→GPU roundtrip overhead until WGSL shaders land).
Sum, Expand, and Cast were pass-through proxies that bypassed the tape. loss = y.Sum() silently produced zero gradients. Now all three record backward ops: SumOp (broadcast grad to input shape), ExpandOp (sum along broadcast dims), CastOp (cast grad back to input dtype). 8 new TDD tests verify gradient flow through each op.
…20 F8) MSELoss.Forward now uses backend.Sum + backend.DivScalar instead of reading GPU data to CPU and summing in a Go loop. All ops recorded on tape, gradients flow correctly: d(MSE)/dx = 2(x-t)/N.
…(ADR-020 F3) RawTensor.Clone() now calls AddRef() through BackendDataRefCounter interface — GPU-agnostic (ADR-019). LazyGPUData.ForceRelease() added for unconditional buffer release. ReclaimMemory removes dead RefCount>1 guard and uses ForceRelease for all non-persistent entries.
…20 F7) applyRotation rewritten with Chunk, Mul, Sub, Add, Cat — all ops go through autodiff tape. Gradients now flow through RoPE, enabling transformer training. ForceNonUnique guards against CPU in-place mutation of shared Chunk buffers.
R1: clearInputBufferCache uses DeferReleaseGPUBuffer instead of immediate Release — prevents GPU panic when called from optimizer.Step() with pending commands in activeBatch. R2: SetTensor takes ownership without Clone/AddRef — prevents GPU memory leak where shared LazyGPUData refcount grew by +params every step. R3: materializeForCPU reads GPU buffer via ReadGPUBuffer without calling Realize() — source tensor stays valid for subsequent backward ops. F5: Backward panics with descriptive message when loss tensor is not found on the tape (was silent fallback to last op). gofmt: cast.go formatted. Fable revalidation tests added (TestRV_*).
kolkov
force-pushed
the
fix/fable-gpu-training-bugs
branch
from
September 25, 2026 16:02
a47bc6b to
401f35b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Critical GPU training bug fixes from Fable 5.1 independent audit (ADR-020). 9 findings fixed, 3 regressions caught and resolved during revalidation.
40 files changed, +1640 / -149 lines. 9 commits.
Autodiff correctness
loss = y.Sum()now produces correct gradients (F6)backend.Sum+backend.DivScalarinstead of CPU readback loop (F8)GPU memory & lifecycle
DeferReleaseGPUBufferprevents panic with pending GPU commands (R1)BackendDataRefCounterinterface,ForceReleaseinReclaimMemory(F3)GPU backward ops
panic("not implemented")(F9)ReadGPUBufferwithoutRealize()side-effect (R3)Other
sync.Map(F4)FeatureSubgroupOperationsTest plan
go test ./internal/...— all packages greenTestGPUTraining_MLP_OOM— 20 steps, pool stable (was panic on R1)TestRV_TrainingLoopMemory— cache[0..0], live[6 6 6 6 6 6 6 6](was[6 8 10 12...]on R2)TestRV_Conv2DBackwardGPU— Conv2D→ReLU→Conv2D 3 steps (was stack overflow on R3)TestRV_DiamondGraph— grad3+exp(a)correctTestBackward_RootsAtLossTensor— metric after loss ignoredTestSum_GradientFlows/TestExpand_GradientFlows/TestCast_GradientFlowsTestRoPE_GradientFlows_3D/_4DTestMSELoss_GradientFlowsgofmt -l .clean