Skip to content

Join named tuples - #161

Merged
Quafadas merged 3 commits into
mainfrom
join-named-tuples
Sep 30, 2026
Merged

Quafadas merged 3 commits into
mainfrom
join-named-tuples

Conversation

@Quafadas

Copy link
Copy Markdown
Owner

No description provided.

Simon Parten and others added 3 commits September 30, 2026 13:14
Joining is the one relational operation that cannot be delegated to the
standard library: stdlib can group and fold rows, but it has no way to
compute the *type* of two named tuples concatenated on a key. Until now
the only cross-table join available was via the scalasql integration,
i.e. only if the data was already in a database.

The key is written after the right hand table - `orders.join(customers)["custId"]`
- because clause interleaving requires a term clause between two type
parameter lists, and an extension's receiver does not count. This is the
same shape stdlib's own `NamedTuple.++` uses.

The left side streams; the right is read into a hash index, lazily, so
building a join drains neither side. A column name appearing on both
sides is a compile error naming the offender, rather than a silently
duplicated column. `leftJoin` optionalises the right hand columns, and
does so idempotently - an already-optional column does not nest - with
`optionalise` as the runtime counterpart of `Optional`.

Only inner and left joins are built in; a right join is a left join with
the tables swapped. Composite keys are not supported.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Treating None as an ordinary value meant missing data squared itself: a
join with three None keys on each side emitted nine rows carrying no
information, and a 10k row join 30% missing in its key produced millions
of junk rows. The key column is exactly where a duplicated value is
least likely to mean "these rows belong together".

So None is now "unknown", and two unknowns are not a match. This is what
SQL does, where NULL = NULL is never true, and what pandas does, dropping
missing keys from a merge. leftJoin still keeps such a row, with its
right hand columns all None.

To match missing to missing deliberately, map the key to a sentinel
first: mapColumn["k", Int](_.getOrElse(-1)).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
After an inner join a None key cannot have matched - the row would have
been dropped - so an Option key is provably present and the Option is
noise. join on custId: Option[Int] now hands back custId: Int, sparing
callers a .get that could never have thrown.

leftJoin deliberately does not narrow: an unmatched left row survives
still holding its None, so Option is the honest type there.

Which leaves the two joins as mirror images. leftJoin ADDS optionality
to the right hand columns, because an unmatched row appears and those
columns really are absent. join REMOVES it from the key, because an
unmatched row does not appear at all - and so the right hand columns
keep their own types, never Option.

Unwrapped is the identity on a non-optional key, so this is a no-op for
the ordinary case, and isOptionKey keeps the runtime unwrap exactly in
step with NarrowKey.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Quafadas
Quafadas merged commit 5b669cf into main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant