Skip to content

fix: critical GPU training bugs from Fable 5.1 audit (ADR-020) - #156

Merged
kolkov merged 13 commits into
mainfrom
fix/fable-gpu-training-bugs
Sep 25, 2026
Merged

kolkov merged 13 commits into
mainfrom
fix/fable-gpu-training-bugs

Conversation

@kolkov

@kolkov kolkov commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Summary

Critical GPU training bug fixes from Fable 5.1 independent audit (ADR-020). 9 findings fixed, 3 regressions caught and resolved during revalidation.

40 files changed, +1640 / -149 lines. 9 commits.

Autodiff correctness

  • Backward roots at loss tensor — was using last tape op, any op after loss corrupted gradients (F5)
  • Sum, Expand, Cast on tape — 3 new backward ops; loss = y.Sum() now produces correct gradients (F6)
  • RoPE gradient flow — rewritten with tensor ops (Chunk/Mul/Sub/Add/Cat), all on tape (F7)
  • MSELoss gradient flow — backend.Sum + backend.DivScalar instead of CPU readback loop (F8)

GPU memory & lifecycle

  • AutodiffBackend forwards ClearInputBufferCache — optimizer cache invalidation reaches GPU backend (F1)
  • Deferred buffer release — DeferReleaseGPUBuffer prevents panic with pending GPU commands (R1)
  • SetTensor ownership — no Clone/AddRef leak, live GPU tensors constant per step (R2)
  • Clone+AddRef — BackendDataRefCounter interface, ForceRelease in ReclaimMemory (F3)
  • ReleaseGradients — public API now accepts backend parameter (F2)

GPU backward ops

  • Conv2D/MaxPool2D backward — CPU fallback instead of panic("not implemented") (F9)
  • materializeForCPU — ReadGPUBuffer without Realize() side-effect (R3)

Other

  • One-hot identity cache per (numClasses, dtype) via sync.Map (F4)
  • Backward panics when loss tensor not on tape (F5 hardening)
  • Subgroup shader tests skip without FeatureSubgroupOperations
  • gogpu deps: wgpu v0.34.4, gputypes v0.8.0, naga v0.19.0

Test plan

  • go test ./internal/... — all packages green
  • TestGPUTraining_MLP_OOM — 20 steps, pool stable (was panic on R1)
  • TestRV_TrainingLoopMemory — cache [0..0], live [6 6 6 6 6 6 6 6] (was [6 8 10 12...] on R2)
  • TestRV_Conv2DBackwardGPU — Conv2D→ReLU→Conv2D 3 steps (was stack overflow on R3)
  • TestRV_DiamondGraph — grad 3+exp(a) correct
  • TestBackward_RootsAtLossTensor — metric after loss ignored
  • TestSum_GradientFlows / TestExpand_GradientFlows / TestCast_GradientFlows
  • TestRoPE_GradientFlows_3D / _4D
  • TestMSELoss_GradientFlows
  • gofmt -l . clean
  • Fable 5.1 revalidation: all blockers closed, PR approved

- AutodiffBackend forwards ClearInputBufferCache to inner backend,
  fixing silent stale-weight training when optimizer runs through
  autodiff wrapper (TASK-154)
- ReclaimMemory clears input buffer cache as defense-in-depth
- Public autodiff.ReleaseGradients accepts optional backend parameter,
  fixing no-op that never released GPU gradient buffers (TASK-155)
- Cache buildOneHotIdentity per (numClasses, dtype) via sync.Map,
  preventing 4 GB/step allocation for large vocab backward (TASK-156)
- Update gogpu deps: wgpu v0.34.4, gputypes v0.8.0, naga v0.19.0
- Fix wgpu API: AdapterInfo/Backends from gputypes, software.NewBackend
tape.Backward now accepts outputTensor parameter to identify the loss
operation on the tape. Backward walks from that op, ignoring any ops
recorded after loss (metrics, logging). Falls back to lastOp when
outputTensor is nil (backward-compatible).
… (ADR-020 F9)

Conv2DInputBackward, Conv2DKernelBackward, and MaxPool2DBackward no
longer panic on WebGPU backend. They materialize GPU tensors to CPU,
delegate to cpu.New(), and return the result. CNN training on GPU now
works (with GPU→CPU→GPU roundtrip overhead until WGSL shaders land).
Sum, Expand, and Cast were pass-through proxies that bypassed the tape.
loss = y.Sum() silently produced zero gradients. Now all three record
backward ops: SumOp (broadcast grad to input shape), ExpandOp (sum
along broadcast dims), CastOp (cast grad back to input dtype).

8 new TDD tests verify gradient flow through each op.
…20 F8)

MSELoss.Forward now uses backend.Sum + backend.DivScalar instead of
reading GPU data to CPU and summing in a Go loop. All ops recorded on
tape, gradients flow correctly: d(MSE)/dx = 2(x-t)/N.
…(ADR-020 F3)

RawTensor.Clone() now calls AddRef() through BackendDataRefCounter
interface — GPU-agnostic (ADR-019). LazyGPUData.ForceRelease() added
for unconditional buffer release. ReclaimMemory removes dead RefCount>1
guard and uses ForceRelease for all non-persistent entries.
…20 F7)

applyRotation rewritten with Chunk, Mul, Sub, Add, Cat — all ops go
through autodiff tape. Gradients now flow through RoPE, enabling
transformer training. ForceNonUnique guards against CPU in-place
mutation of shared Chunk buffers.
R1: clearInputBufferCache uses DeferReleaseGPUBuffer instead of
immediate Release — prevents GPU panic when called from optimizer.Step()
with pending commands in activeBatch.

R2: SetTensor takes ownership without Clone/AddRef — prevents GPU memory
leak where shared LazyGPUData refcount grew by +params every step.

R3: materializeForCPU reads GPU buffer via ReadGPUBuffer without
calling Realize() — source tensor stays valid for subsequent backward ops.

F5: Backward panics with descriptive message when loss tensor is not
found on the tape (was silent fallback to last op).

gofmt: cast.go formatted.

Fable revalidation tests added (TestRV_*).
@kolkov
kolkov force-pushed the fix/fable-gpu-training-bugs branch from a47bc6b to 401f35b Compare September 25, 2026 16:02
@kolkov
kolkov merged commit cc3d257 into main Sep 25, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant