Skip to content

test: one model, one snippet, two profiles, two forms, both correct - #1125

Merged
michalharakal merged 1 commit into
developfrom
feature/1118-weight-form-acceptance
Aug 25, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1118-weight-form-acceptance

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1118. Slice 5 of 5 for #1109.

The acceptance criterion for the whole thing: a model author writes nothing about weight forms, and the same code is correct on two very different devices.

Everything else in #1109 is machinery for that one claim, so the test is arranged to make it falsifiable rather than to exercise the machinery:

private fun userCode(x: Tensor<FP32, Float>, w: Tensor<FP32, Float>): FloatArray =
    x.matmulWeightTransposed(w).data.copyToFloatArray()

No profile, no form, no policy. Called identically on both paths. If honouring a device ever required the caller to say something, this file would have to change to keep passing.

Desktop resolves to KeepAsStored + HEAP; the 2 GB board to DequantizeTo(FP32) + MAPPED. Asserted to be different, then asserted to agree numerically within quantization error — not bit-identical, and the test says why: one path multiplies through a Q8_0 kernel and the other through dequantized floats.

It found two problems before asserting anything

Slices 1 and 2 disagreed. The resolver asked for KERNEL_FEED in the common case — a target that has kernels — and the loader rejected it pending #1120. Neither slice's own tests could see this; only running them together could.

The fix is that wanting feed order and being able to produce it are facts about different components. KernelCapabilities.wantsKernelFeedOrder answers the kernel's half; the new canProduceKernelFeedOrder parameter answers the loader's, and is false until #1120. Collapsing them into one is what let slice 1 hand slice 2 a form it had to reject.

A pre-existing silent correctness bug — #1124. An output mismatch of 12.94 vs -71.34 was too large for quantization error. The weight's own toFloatArray() agreed with the dense path and not with the packed one, which narrowed it to: the common DefaultCpuOps computes packed matmul wrongly when the KernelRegistry is empty.

ops providers result
DefaultCpuOpsJvm none / registered correct
DefaultCpuOps (common) scalar registered correct
DefaultCpuOps (common) none wrong

Not caused by this work, and not fixed here. The test registers ScalarKernelProvider — what every PlatformCpuOpsFactory does at startup — so it runs in the configuration a real application runs in, with a comment saying so and pointing at #1124.

The other two tests

The mobile plan knows what its form costs before the load (#1116's pricing, end to end), and a strict board is told about a missing kernel instead of quietly paying 4× for it — MOBILE_2GB's own documentation calls dispatcher-inserted dequantization "the defect it is".

Build change

skainet-io-gguf's jvmTest gains skainet-backend-cpu and skainet-backend-api. Test-only, and not a cycle — the CPU backend does not know about GGUF. This module's own DefaultDataExecutionContext carries VoidTensorOps, and an acceptance test that loads a model has to actually run it.

Gate

scripts/pr-gate.sh — all legs passed.

🤖 Generated with Claude Code

Closes #1118. Slice 5 of #1109.

The acceptance criterion for the whole of #1109: a model author writes
nothing about weight forms, and the same code is correct on a workstation and
on a 2 GB board because the resolver answered differently. The test is
arranged so that claim is falsifiable — `userCode` takes a weight and an
activation, no profile, no form, no policy, and is called identically on both
paths. If honouring a device ever required the caller to say something, this
file would have to change to keep passing.

It found two things before it asserted anything.

Slices 1 and 2 disagreed. The resolver asked for KERNEL_FEED in the common
case of a target that has kernels, and the loader rejected it pending #1120 —
neither slice's own tests could see this, only running them together could.
Wanting feed order is a fact about the kernel; producing it is a fact about
whoever writes the bytes, and collapsing the two is what let slice 1 emit a
form nothing could honour. The resolver now takes
canProduceKernelFeedOrder, false until #1120 makes it true.

And a pre-existing silent correctness bug, filed as #1124: the common
DefaultCpuOps computes packed matmul wrongly when the KernelRegistry is
empty. Chasing an output mismatch of 12.94 vs -71.34 established that the
weight's own toFloatArray() agreed with the dense path and not with the
packed one. The test registers ScalarKernelProvider, as every
PlatformCpuOpsFactory does at startup, so it runs in the configuration a real
application runs in rather than in the broken one.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1125 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit cff2c7f into develop Aug 25, 2026
18 checks passed
@michalharakal
michalharakal deleted the feature/1118-weight-form-acceptance branch August 25, 2026 15:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Same model, same code, two profiles, two forms, both run

1 participant