Repository navigation
Conversation
aunjgr
marked this pull request as ready for review
October 8, 2026 02:48
1 task done
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Add the native coefficient, codec and scalar-kernel foundation for MatrixOne's approved exact-decimal contract. Preserve independent physical width/precision/scale, signed little-endian Decimal64/128/256 coefficients, strict NULL and execution masks, once-rounded arithmetic, MO-bound division scales and operation-owned scalar/cast errors.
Decimal256 uses a parent-nullable STRUCT with a signed high limb and three unsigned limbs. Its input/result codecs retain existing admission ownership; result interleaving reuses admitted scratch and visits output blocks once per chunk. Checked kernels use the caller's stream and resource, retain unsafe owners on unprovable stream quiescence, and use 64-bit grid-stride counters. Narrow cuDF paths are used only for complete domains that cannot overflow; widening casts are required by the pinned cuDF and are retained.
This is implementation PR 3 in matrixorigin/matrixone#28966's approved series. Approved numeric design: matrixorigin/matrixone#29449, exact document blob
42a89f09a1d168d02b9583cb3ea7b4de6dbb5634. Base isupstream-dev-mergeate2e2f08f9fd1eaa1253297493df0ab11a6078664. Pin the importer to merged matrixorigin/duckdb-substrait#3 (dd4cab14b82754ca919633913436cd530f496e00), without unrelated dependency synchronization. Its unity-test CI correction is separate in matrixorigin/duckdb-substrait#4.The increment does not register the exact-decimal Substrait family, advertise its capability, admit additional MO queries or change ordinary DuckDB decimal execution. Aggregates/keys, numeric plan integration, typed public statuses, MO lowering and the public all-22 campaign remain subsequent deliverables. ABI-v1 layouts and existing capability bits are preserved. No default/cutover or Flight removal occurs here.
Validation
Frozen Sirius
moPixi environment, incremental host build, DuckDB069cc9f9b5be802405797faecc284961b07c70ef, CUDA 13.3.73, GCC 14.4.0, RTX 3070 with driver 615.71.09. No image rebuild.Fractionarithmetic.__int128oracle, scalar versus cast errors, inactive/NULL controls, sliced operands, bitmap warp/tail boundaries, Decimal256 chunked/sliced output, and existing constant/NULL/retry decoding and 1/2/4-worker cases.docs/mo-exact-decimal-primitives.md; its target is excluded from default builds. Raw repetitions and provenance are retained with the implementation.The local sample uses 262,144 rows, two warm-ups and seven measured repetitions, with one 64 MiB initial / 128 MiB maximum caller-owned pool and one task stream. Median wall times: unchecked cuDF add64 0.195870 ms; checked add64 0.027580 ms; equivalent widened cuDF multiply64-to128 0.499538 ms; checked multiply64-to128 0.500228 ms; checked add256 0.262309 ms; checked divide256 4.662440 ms. The multiply overhead in this sample is about 0.14%. Decimal256 has no equivalent cuDF carrier; these are kernel samples, not all-22 query timing or rollout acceptance.
Self-review
Ownership audit: Q1 closes through returned owners or existing fatal process quarantine; Q2 uses the caller's task stream and existing fatal-health path when quiescence fails; Q3 charges row buffers to the caller resource and codec storage to existing windows. No new thread, global registry, pool or query lifecycle is introduced in production. The benchmark owns its separate bounded pool.
No new public SQL path is enabled, so this increment's proof is native UT/GPU consumer evidence. Public SQL/error/metadata parity, full numeric capability, SF1/SF10 Q1-Q22 at streams=2 plus 1/4 controls, paused-consumer/cancellation/async-failure/memory-baseline gates and numeric-heavy query performance remain mandatory before cutover.
Checklist
References
Refs matrixorigin/matrixone#28968 and matrixorigin/matrixone#28966. Does not close either issue.