Conversation
…lary Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ventNgramLookupLayer
| if all(id_value == 0 for id_value in event_tuple): | ||
| all_tokens.extend(pad_tokens) | ||
| else: | ||
| all_tokens.extend(lookup_table.get(event_tuple, unk_tokens)) |
There was a problem hiding this comment.
Just to double check, is it intended that unseen items become <unk> even when their n-grams are in the vocab? For my understanding, e.g., a new (1, 10, 999) would lose its L0=1 / L1=10 tokens
There was a problem hiding this comment.
Yes, that's intended. The fitted artefact maps a whole-tuple to tokens: each tuple's tokens are computed once at fit time, so the transform and the Keras layer are a single static lookup per event (tuple).
We could change to n-gram level instead, I just made this convention to keep serving to one lookup per event. Handling unknowns per n-gram would mean forming up to 2^tupleSize keys per event in the graph, looking each one up, and selecting the topK.
Still static and not really an overhead, I just didn't consider it in scope for this PR, and would have to carry out full end2end experiments to compare performance, or add it in a backward-compatible manner that I haven't tested yet.
There was a problem hiding this comment.
Gotcha thanks, then this PR looks good as is 👌
Description
Adds a tokenizer for sequences of fixed-size discrete-ID tuples.
An event is one such tuple:
tupleSizeIDs (default 4) that together describe one thing, for example one item's 4-level semantic ID(L0, L1, L2, L3). An input column holds a flat integer array of one or more events laid end to end. A user's last 10 interacted items with 4-level IDs, for instance, is one row of10 × 4 = 40integers;numEventsPerInputgives each column's event count, and all-zero events are padding.The estimator learns a shared n-gram vocabulary over the corpus, and each event is turned into a handful of integer tokens, ready for one shared embedding table.
Classes
EventNgramLookupEstimator: counts within-event n-grams, keeps thevocabSizemost frequent aboveminNgramFreq, and pre-computes each tuple'stopKtokens.EventNgramLookupTransformer: applies the fitted table in Spark.EventNgramLookupLayer: applies the same table in-graph at serving. One layer handles all input columns, using a singleIntegerLookupon packedint64keys, so there are no string ops.How it behaves
(L0, L2)is an n-gram.0is padding and token1is unknown; learned n-grams get ids from2. On the input side, an ID of0marks an absent level and an all-zero event is padding. Any other event without a learned token becomes a single<unk>, whether it was never seen during fitting or matched nothing.includeTokenTypesadds a parallel<col>_typescolumn with each token's ID-level bitmask.Testing
Unit and Spark vs Keras parity tests (rank-2 and rank-3, several tuple sizes); also validated end to end in a downstream pipeline on production-scale data, including a SavedModel export.
Keras Layer Checklist
Verify that:
_callmethod has been implemented in the new layer.compatible_dtypesproperty is defined in the new layer.@tf.keras.utils.register_keras_serializable(package=kamae.__name__).name,input_dtype, andoutput_dtypeas arguments to the constructor and that this is passed to the super constructor.get_configmethod.layersdirectory.Spark Transformer/Estimator Checklist
Verify that:
__init__andsetParamsmethods.Paramsclass here.compatible_dtypesproperty has been implemented to specify the input/output data types that my transformer/estimator supports.get_tf_layermethod.transformers/estimatorsdirectory.Finally, please verify that: