Skip to content

feat(numeric): add checked MO exact-decimal primitives - #25

Merged
aunjgr merged 1 commit into
matrixorigin:upstream-dev-mergefrom
aunjgr:feature/28968-mo-decimal-primitives
Oct 8, 2026
Merged

aunjgr merged 1 commit into
matrixorigin:upstream-dev-mergefrom
aunjgr:feature/28968-mo-decimal-primitives

Conversation

@aunjgr

@aunjgr aunjgr commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Description

Add the native coefficient, codec and scalar-kernel foundation for MatrixOne's approved exact-decimal contract. Preserve independent physical width/precision/scale, signed little-endian Decimal64/128/256 coefficients, strict NULL and execution masks, once-rounded arithmetic, MO-bound division scales and operation-owned scalar/cast errors.

Decimal256 uses a parent-nullable STRUCT with a signed high limb and three unsigned limbs. Its input/result codecs retain existing admission ownership; result interleaving reuses admitted scratch and visits output blocks once per chunk. Checked kernels use the caller's stream and resource, retain unsafe owners on unprovable stream quiescence, and use 64-bit grid-stride counters. Narrow cuDF paths are used only for complete domains that cannot overflow; widening casts are required by the pinned cuDF and are retained.

This is implementation PR 3 in matrixorigin/matrixone#28966's approved series. Approved numeric design: matrixorigin/matrixone#29449, exact document blob 42a89f09a1d168d02b9583cb3ea7b4de6dbb5634. Base is upstream-dev-merge at e2e2f08f9fd1eaa1253297493df0ab11a6078664. Pin the importer to merged matrixorigin/duckdb-substrait#3 (dd4cab14b82754ca919633913436cd530f496e00), without unrelated dependency synchronization. Its unity-test CI correction is separate in matrixorigin/duckdb-substrait#4.

The increment does not register the exact-decimal Substrait family, advertise its capability, admit additional MO queries or change ordinary DuckDB decimal execution. Aggregates/keys, numeric plan integration, typed public statuses, MO lowering and the public all-22 campaign remain subsequent deliverables. ABI-v1 layouts and existing capability bits are preserved. No default/cutover or Flight removal occurs here.

Validation

Frozen Sirius mo Pixi environment, incremental host build, DuckDB 069cc9f9b5be802405797faecc284961b07c70ef, CUDA 13.3.73, GCC 14.4.0, RTX 3070 with driver 615.71.09. No image rebuild.

  • CPU arithmetic: 52,330 assertions / 5 cases pass, including 96 full-width vectors independently rechecked with Python Fraction arithmetic.
  • Native control/input/result ownership suite: 594 assertions / 77 cases pass.
  • GPU arithmetic and affected native input/result cases: 29,795 assertions / 11 cases pass. Includes signed physical boundaries against an independent host __int128 oracle, scalar versus cast errors, inactive/NULL controls, sliced operands, bitmap warp/tail boundaries, Decimal256 chunked/sliced output, and existing constant/NULL/retry decoding and 1/2/4-worker cases.
  • Pinned pre-commit static/formatting hooks and orphan-test registration pass. The new tests are wired into independently runnable native targets.
  • Reproducible standalone benchmark and commands are in docs/mo-exact-decimal-primitives.md; its target is excluded from default builds. Raw repetitions and provenance are retained with the implementation.

The local sample uses 262,144 rows, two warm-ups and seven measured repetitions, with one 64 MiB initial / 128 MiB maximum caller-owned pool and one task stream. Median wall times: unchecked cuDF add64 0.195870 ms; checked add64 0.027580 ms; equivalent widened cuDF multiply64-to128 0.499538 ms; checked multiply64-to128 0.500228 ms; checked add256 0.262309 ms; checked divide256 4.662440 ms. The multiply overhead in this sample is about 0.14%. Decimal256 has no equivalent cuDF carrier; these are kernel samples, not all-22 query timing or rollout acceptance.

Self-review

Closure Risk and ownership Evidence / terminal behavior
Fixed-width arithmetic and descriptors R2; no per-row allocation; 512-bit local scratch; exact declared/physical domains CPU and independent rational/integer oracles; GPU masks/errors/boundaries
GPU values, validity and error storage R3; caller resource/stream; result owns buffers; caller retains operands Quiescent success/error; fatal quarantine reuses existing admission-sealing process owner; GPU consumer tests
Decimal256 input/result codecs R3; existing 64 MiB windows, leases and admitted scratch Flat/constant/NULL/retry input; segmented storage, chunk boundary, parent slice plus nonzero begin; no complete-result copy
Dependency/build/docs/benchmark R0/R1; merged importer pin; explicit test/benchmark targets Static hooks, build, nonempty selected tests and seven raw samples

Ownership audit: Q1 closes through returned owners or existing fatal process quarantine; Q2 uses the caller's task stream and existing fatal-health path when quiescence fails; Q3 charges row buffers to the caller resource and codec storage to existing windows. No new thread, global registry, pool or query lifecycle is introduced in production. The benchmark owns its separate bounded pool.

No new public SQL path is enabled, so this increment's proof is native UT/GPU consumer evidence. Public SQL/error/metadata parity, full numeric capability, SF1/SF10 Q1-Q22 at streams=2 plus 1/4 controls, paused-consumer/cancellation/async-failure/memory-baseline gates and numeric-heavy query performance remain mandatory before cutover.

Checklist

  • Read CONTRIBUTING.md and meet the PR reviewability checklist
  • Cover changes with focused existing/new tests
  • Document the native contract and reproducible validation commands
  • No production configuration/default change

References

Refs matrixorigin/matrixone#28968 and matrixorigin/matrixone#28966. Does not close either issue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant