Skip to content

feat: add EventNgramLookup estimator, transformer and layer - #77

Open
Sbranikas wants to merge 7 commits into
mainfrom
feature/event-ngram-lookup
Open

Sbranikas wants to merge 7 commits into
mainfrom
feature/event-ngram-lookup

Conversation

@Sbranikas

@Sbranikas Sbranikas commented Oct 1, 2026 •

Copy link
Copy Markdown

Description

Adds a tokenizer for sequences of fixed-size discrete-ID tuples.

An event is one such tuple: tupleSize IDs (default 4) that together describe one thing, for example one item's 4-level semantic ID (L0, L1, L2, L3). An input column holds a flat integer array of one or more events laid end to end. A user's last 10 interacted items with 4-level IDs, for instance, is one row of 10 × 4 = 40 integers; numEventsPerInput gives each column's event count, and all-zero events are padding.

The estimator learns a shared n-gram vocabulary over the corpus, and each event is turned into a handful of integer tokens, ready for one shared embedding table.

Classes

  • EventNgramLookupEstimator: counts within-event n-grams, keeps the vocabSize most frequent above minNgramFreq, and pre-computes each tuple's topK tokens.
  • EventNgramLookupTransformer: applies the fitted table in Spark.
  • EventNgramLookupLayer: applies the same table in-graph at serving. One layer handles all input columns, using a single IntegerLookup on packed int64 keys, so there are no string ops.

How it behaves

  • N-grams stay within an event and can skip levels. For example, (L0, L2) is an n-gram.
  • Reserved token ids in the vocabulary: in the output, token 0 is padding and token 1 is unknown; learned n-grams get ids from 2. On the input side, an ID of 0 marks an absent level and an all-zero event is padding. Any other event without a learned token becomes a single <unk>, whether it was never seen during fitting or matched nothing.
  • Optional types: includeTokenTypes adds a parallel <col>_types column with each token's ID-level bitmask.
  • Spark and Keras agree. They produce the same tokens for every input both accept, and they reject the same inputs: non-negative integer IDs only, as a whole number of events.

Testing

Unit and Spark vs Keras parity tests (rank-2 and rank-3, several tuple sizes); also validated end to end in a downstream pipeline on production-scale data, including a SavedModel export.

Keras Layer Checklist

Verify that:

  • The new Keras layer extends BaseLayer
  • The _call method has been implemented in the new layer.
  • The compatible_dtypes property is defined in the new layer.
  • The new layer is decorated with @tf.keras.utils.register_keras_serializable(package=kamae.__name__).
  • The new layer takes a name, input_dtype, and output_dtype as arguments to the constructor and that this is passed to the super constructor.
  • The Keras layer is serializable. I have implemented the get_config method.
  • There are unit tests of the new layer.
  • There is a specific test of layer serialisation added here.
  • The new layer is imported in the init.py file in the layers directory.

Spark Transformer/Estimator Checklist

Verify that:

  • The new Spark Transformer extends BaseTransformer.
  • If the new transform needs a fit method, a Spark Estimator has been implemented that extends BaseEstimator.
  • The instructions in the above docs page have been followed for the __init__ and setParams methods.
  • The transformer uses one of the input/output mixin classes from base.py.
  • If the new transformer requires more parameters that would need to be serialised to the Spark ML pipeline, there is a implemented parameter class by extending the Params class here.
  • The compatible_dtypes property has been implemented to specify the input/output data types that my transformer/estimator supports.
  • A Keras subclassed layer is returned in the transformer's get_tf_layer method.
  • There are unit tests of the new transform. In particular, there are parity tests between the Spark and Keras implementations.
  • The new transformer/estimator is imported in the init.py file in the transformers/estimators directory.

Finally, please verify that:

  • There is a new entry (alphabetical order) in the README table describing the new layer/transformer

@Sbranikas
Sbranikas requested a review from a team as a code owner October 1, 2026 08:55
if all(id_value == 0 for id_value in event_tuple):
all_tokens.extend(pad_tokens)
else:
all_tokens.extend(lookup_table.get(event_tuple, unk_tokens))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to double check, is it intended that unseen items become <unk> even when their n-grams are in the vocab? For my understanding, e.g., a new (1, 10, 999) would lose its L0=1 / L1=10 tokens

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, that's intended. The fitted artefact maps a whole-tuple to tokens: each tuple's tokens are computed once at fit time, so the transform and the Keras layer are a single static lookup per event (tuple).

We could change to n-gram level instead, I just made this convention to keep serving to one lookup per event. Handling unknowns per n-gram would mean forming up to 2^tupleSize keys per event in the graph, looking each one up, and selecting the topK.
Still static and not really an overhead, I just didn't consider it in scope for this PR, and would have to carry out full end2end experiments to compare performance, or add it in a backward-compatible manner that I haven't tested yet.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Gotcha thanks, then this PR looks good as is 👌

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants