Skip to content

Atomnorm#2209

Open
ArthurZucker wants to merge 34 commits into
feat/train_encode_splitfrom
atomnorm
Open

Atomnorm#2209
ArthurZucker wants to merge 34 commits into
feat/train_encode_splitfrom
atomnorm

Conversation

@ArthurZucker

@ArthurZucker ArthurZucker commented Jul 16, 2026

Copy link
Copy Markdown
Collaborator

PipelineTokenizer benchmark

7 / 8 models supported — PipelineTokenizer vs tokenizers v0.23.1 (latest release) · ~10 kB inputs · single thread + 1/2/4/8/max-thread sweep

141e001f2 · 2026-07-16 12:13 UTC · Intel(R) Xeon(R) Platinum 8375C CPU @ 2.90GHz · 32 cores

Per-model encode throughput vs latest release

vs base branch (2c4f6100b) — per-model geomean ×speedup of this PR's PipelineTokenizer against the base branch's; regressions in red.

Per-model encode throughput vs base branch Per-model memory footprint Minimal encode binary size
bert-base-uncased — normalizer-heavy WordPiece · ×15.76 vs v0.23.1 · ×2.81 vs base bert-base-uncased speedup bert-base-uncased stage decomposition bert-base-uncased thread scaling

Memory (RSS MB, load+encode): v0.23.1 2+0 (peak 6) · Pipeline 7+0 (peak 7)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 7.1 94.6 ×13.41 ×3.44 12% (1.2) 24% (2.5) 39% (4.0) 26% (2.6) match
arb_Arab lang 3.5 63.7 ×18.45 ×2.55 8% (1.2) 25% (3.8) 17% (2.6) 50% (7.7) match
ben_Beng lang 5.2 81.0 ×15.64 ×2.32 10% (1.2) 27% (3.2) 21% (2.5) 43% (5.2) match
cmn_Hani lang 3.1 51.1 ×16.37 ×2.71 6% (1.2) 26% (4.9) 24% (4.6) 44% (8.5) match
ell_Grek lang 3.2 57.1 ×17.71 ×2.40 7% (1.2) 29% (4.9) 15% (2.6) 49% (8.4) match
eng_Latn lang 3.8 58.6 ×15.23 ×3.20 15% (2.5) 13% (2.1) 21% (3.4) 52% (8.6) match
heb_Hebr lang 3.5 67.5 ×19.52 ×3.39 8% (1.2) 23% (3.3) 19% (2.7) 51% (7.3) match
hin_Deva lang 5.6 81.3 ×14.49 ×2.96 10% (1.2) 30% (3.6) 21% (2.5) 40% (4.8) match
jpn_Jpan lang 3.7 56.5 ×15.38 ×2.06 7% (1.2) 28% (4.9) 24% (4.1) 41% (7.1) match
kat_Geor lang 5.8 94.8 ×16.44 ×3.37 11% (1.2) 23% (2.4) 23% (2.4) 43% (4.4) match
kor_Hang lang 2.1 31.8 ×14.85 ×1.80 4% (1.2) 35% (11.0) 21% (6.4) 40% (12.5) match
rus_Cyrl lang 3.1 57.0 ×18.17 ×2.41 7% (1.2) 21% (3.5) 14% (2.5) 58% (9.9) match
tam_Taml lang 6.5 96.0 ×14.68 ×2.51 11% (1.2) 34% (3.4) 20% (2.1) 35% (3.5) match
tha_Thai lang 8.2 126.7 ×15.52 ×3.83 15% (1.2) 45% (3.4) 20% (1.5) 20% (1.5) match
added_normalized_dense modalities 6.2 76.4 ×12.39 ×3.92 11% (1.3) 32% (3.9) 50% (6.1) 7% (0.9) match
added_normalized_sparse modalities 4.8 61.1 ×12.62 ×3.35 12% (1.9) 21% (3.3) 37% (5.8) 29% (4.5) match
added_special_dense modalities 4.8 60.0 ×12.38 ×1.57 32% (5.0) 8% (1.3) 48% (7.4) 11% (1.8) match
added_special_sparse modalities 3.7 50.2 ×13.55 ×2.37 20% (3.7) 13% (2.5) 36% (6.8) 31% (5.9) match
agentic-traces modalities 3.3 56.5 ×17.10 ×3.11 14% (2.3) 12% (2.1) 22% (3.9) 52% (8.9) match
agentic_swe modalities 3.5 72.4 ×20.48 ×3.72 14% (1.9) 16% (2.1) 19% (2.6) 51% (6.8) match
code_mixed modalities 3.4 65.6 ×19.38 ×3.46 14% (2.1) 14% (2.1) 22% (3.2) 50% (7.4) match
math_latex modalities 3.4 57.3 ×16.64 ×3.15 14% (2.4) 13% (2.1) 22% (3.7) 51% (8.8) match
deepseek-v4 — deepseek 3-regex split-heavy byte-level BPE · ×5.19 vs v0.23.1 · ×1.01 vs base deepseek-v4 speedup deepseek-v4 stage decomposition deepseek-v4 thread scaling

Memory (RSS MB, load+encode): v0.23.1 52+0 (peak 58) · Pipeline 85+0 (peak 85)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 3.9 33.8 ×8.57 ×1.00 2% (0.7) 0% (0.0) 16% (4.7) 81% (23.2) match
arb_Arab lang 3.7 16.6 ×4.54 ×0.98 1% (0.6) 0% (0.0) 6% (3.4) 93% (54.9) match
ben_Beng lang 4.5 18.4 ×4.08 ×1.02 1% (0.6) 0% (0.0) 6% (3.1) 93% (50.5) match
cmn_Hani lang 3.4 16.7 ×4.87 ×1.01 1% (0.8) 0% (0.0) 6% (3.2) 93% (52.6) match
ell_Grek lang 3.7 18.0 ×4.83 ×0.99 1% (0.6) 0% (0.0) 6% (3.4) 93% (50.7) match
eng_Latn lang 2.5 12.6 ×5.11 ×0.99 3% (2.0) 0% (0.0) 7% (5.2) 91% (70.6) match
heb_Hebr lang 3.2 12.8 ×4.02 ×1.01 1% (0.6) 0% (0.0) 5% (3.6) 95% (72.6) match
hin_Deva lang 4.1 22.0 ×5.36 ×1.01 1% (0.6) 0% (0.0) 7% (3.3) 91% (40.8) match
jpn_Jpan lang 3.8 18.3 ×4.77 ×1.00 1% (0.7) 0% (0.0) 6% (3.0) 93% (49.4) match
kat_Geor lang 4.5 18.3 ×4.05 ×1.03 1% (0.6) 0% (0.0) 5% (2.9) 93% (49.9) match
kor_Hang lang 3.4 19.7 ×5.79 ×0.99 1% (0.6) 0% (0.0) 7% (3.7) 91% (44.9) match
rus_Cyrl lang 3.7 14.7 ×3.98 ×1.02 1% (0.6) 0% (0.0) 5% (3.3) 94% (62.9) match
tam_Taml lang 4.8 17.8 ×3.69 ×1.03 1% (0.6) 0% (0.0) 5% (2.7) 94% (52.2) match
tha_Thai lang 5.6 14.9 ×2.67 ×1.00 1% (0.6) 0% (0.0) 3% (2.2) 96% (63.3) match
added_normalized_dense modalities 4.0 23.6 ×5.97 ×1.03 2% (0.8) 0% (0.0) 6% (2.7) 92% (37.5) match
added_normalized_sparse modalities 3.6 19.7 ×5.41 ×1.03 3% (1.3) 0% (0.0) 8% (3.8) 90% (44.5) match
added_special_dense modalities 2.9 35.9 ×12.27 ×0.99 25% (6.6) 1% (0.4) 22% (5.8) 51% (13.2) match
added_special_sparse modalities 3.2 20.0 ×6.29 ×0.99 8% (3.7) 1% (0.3) 13% (6.5) 79% (38.5) match
agentic-traces modalities 2.4 14.0 ×5.80 ×1.00 3% (1.8) 0% (0.0) 8% (5.4) 90% (62.9) match
agentic_swe modalities 2.4 15.0 ×6.36 ×1.02 2% (1.3) 0% (0.0) 6% (3.8) 92% (59.8) match
code_mixed modalities 2.7 15.4 ×5.81 ×1.01 2% (1.5) 0% (0.0) 7% (4.6) 91% (62.4) match
math_latex modalities 2.5 13.9 ×5.52 ×1.01 3% (1.9) 0% (0.0) 8% (5.3) 90% (62.2) match

Pre-tokenize: classify + fsm vs regex engines — ns/byte, lower better. The fsm is the scalar jump-table in both pipe columns; SIMD / scalar is the classify pass (regex pre-tokenizers have no SIMD fsm). ×vs = engine ÷ our pipeline (SIMD / scalar classify); onig & pcre2 (JIT) are C, fancy is pure-Rust fancy-regex, logos is a compile-time DFA lexer (approximate grammar; n/a for deepseek).

Fixture classify SIMD classify scalar pipe (SIMD cls + fsm) pipe (scalar cls + fsm) onig fancy pcre2 logos ×vs onig ×vs fancy ×vs pcre2 ×vs logos
amh_Ethi 2.80 3.43 4.69 5.32 41.7 19.8 9.1 8.9× / 7.8× 4.2× / 3.7× 1.9× / 1.7×
arb_Arab 1.15 2.97 3.39 5.21 49.3 22.6 10.1 14.6× / 9.5× 6.7× / 4.3× 3.0× / 1.9×
ben_Beng 1.62 2.95 3.11 4.44 36.6 16.1 7.6 11.8× / 8.2× 5.2× / 3.6× 2.4× / 1.7×
cmn_Hani 1.12 2.22 3.18 4.28 63.0 34.2 14.2 19.8× / 14.7× 10.8× / 8.0× 4.5× / 3.3×
ell_Grek 0.58 2.95 3.44 5.81 48.7 20.6 9.6 14.2× / 8.4× 6.0× / 3.5× 2.8× / 1.7×
eng_Latn 0.09 1.46 5.17 6.54 67.3 39.0 16.2 13.0× / 10.3× 7.5× / 6.0× 3.1× / 2.5×
heb_Hebr 1.15 3.01 3.55 5.41 51.7 23.6 10.5 14.5× / 9.6× 6.6× / 4.4× 2.9× / 1.9×
hin_Deva 1.51 3.08 3.29 4.86 39.0 18.7 8.6 11.9× / 8.0× 5.7× / 3.9× 2.6× / 1.8×
jpn_Jpan 1.70 3.34 3.00 4.63 53.0 27.5 11.8 17.7× / 11.4× 9.2× / 5.9× 3.9× / 2.5×
kat_Geor 1.53 2.57 2.92 3.96 32.6 14.9 7.2 11.1× / 8.2× 5.1× / 3.8× 2.5× / 1.8×
kor_Hang 1.22 2.71 3.67 5.15 51.4 27.1 11.7 14.0× / 10.0× 7.4× / 5.3× 3.2× / 2.3×
rus_Cyrl 1.16 2.93 3.28 5.04 47.5 20.5 9.5 14.5× / 9.4× 6.3× / 4.1× 2.9× / 1.9×
tam_Taml 0.93 2.99 2.71 4.78 32.9 13.4 6.6 12.1× / 6.9× 4.9× / 2.8× 2.4× / 1.4×
tha_Thai 1.51 2.55 2.16 3.20 28.0 10.1 5.5 13.0× / 8.8× 4.7× / 3.2× 2.5× / 1.7×
added_normalized_dense 0.06 1.47 2.66 4.07 43.2 20.0 9.2 16.3× / 10.6× 7.5× / 4.9× 3.5× / 2.3×
added_normalized_sparse 0.06 1.47 3.76 5.18 51.2 25.4 11.5 13.6× / 9.9× 6.8× / 4.9× 3.1× / 2.2×
added_special_dense 0.06 1.77 5.76 7.47 173.2 97.7 38.8 30.1× / 23.2× 17.0× / 13.1× 6.7× / 5.2×
added_special_sparse 0.06 1.47 6.48 7.89 106.9 58.0 24.2 16.5× / 13.5× 9.0× / 7.4× 3.7× / 3.1×
agentic-traces 0.74 1.47 5.40 6.14 83.6 51.6 20.1 15.5× / 13.6× 9.5× / 8.4× 3.7× / 3.3×
agentic_swe 0.67 1.44 3.82 4.59 100.1 64.5 23.5 26.2× / 21.8× 16.9× / 14.0× 6.2× / 5.1×
code_mixed 0.06 1.44 4.61 6.00 75.8 53.0 18.4 16.4× / 12.6× 11.5× / 8.8× 4.0× / 3.1×
math_latex 0.73 1.48 5.34 6.09 80.7 47.2 19.3 15.1× / 13.3× 8.8× / 7.7× 3.6× / 3.2×
gpt2 — gpt2 ByteLevel regex · ×9.80 vs v0.23.1 · ×1.00 vs base gpt2 speedup gpt2 stage decomposition gpt2 thread scaling

Memory (RSS MB, load+encode): v0.23.1 17+1 (peak 19) · Pipeline 20+0 (peak 20)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 4.1 49.7 ×12.20 ×1.02 3% (0.7) 0% (0.0) 19% (3.7) 77% (14.9) match
arb_Arab lang 3.2 25.9 ×8.11 ×0.96 2% (0.6) 0% (0.0) 6% (2.3) 92% (35.0) match
ben_Beng lang 2.3 46.4 ×19.83 ×0.98 3% (0.6) 0% (0.0) 14% (3.0) 83% (17.4) match
cmn_Hani lang 3.3 29.2 ×8.84 ×1.02 2% (0.6) 0% (0.0) 7% (2.3) 91% (30.4) match
ell_Grek lang 3.3 28.7 ×8.77 ×1.01 2% (0.6) 0% (0.0) 7% (2.2) 92% (31.1) match
eng_Latn lang 2.9 13.8 ×4.83 ×0.99 3% (2.0) 0% (0.0) 4% (3.1) 93% (65.9) match
heb_Hebr lang 3.2 29.0 ×9.06 ×0.98 2% (0.6) 0% (0.0) 7% (2.3) 91% (31.1) match
hin_Deva lang 2.6 39.7 ×15.04 ×1.01 2% (0.6) 0% (0.0) 13% (3.1) 85% (20.8) match
jpn_Jpan lang 3.6 21.9 ×6.07 ×1.00 1% (0.6) 0% (0.0) 5% (2.1) 94% (42.0) match
kat_Geor lang 3.7 72.8 ×19.63 ×1.00 5% (0.6) 0% (0.0) 16% (2.1) 79% (10.4) match
kor_Hang lang 3.1 42.6 ×13.80 ×1.01 3% (0.6) 0% (0.0) 11% (2.5) 86% (19.4) match
rus_Cyrl lang 3.3 27.8 ×8.36 ×1.02 2% (0.6) 0% (0.0) 6% (2.1) 92% (32.9) match
tam_Taml lang 2.2 67.2 ×30.38 ×1.00 4% (0.6) 0% (0.0) 19% (2.6) 77% (11.0) match
tha_Thai lang 2.9 35.3 ×12.30 ×0.94 2% (0.6) 0% (0.0) 9% (2.5) 89% (24.4) match
added_normalized_dense modalities 3.6 26.1 ×7.17 ×1.02 2% (0.8) 0% (0.0) 3% (1.0) 95% (34.8) match
added_normalized_sparse modalities 3.4 21.8 ×6.42 ×1.00 3% (1.3) 0% (0.0) 4% (2.0) 93% (41.5) match
added_special_dense modalities 3.5 45.4 ×12.86 ×0.99 24% (4.9) 1% (0.1) 19% (3.8) 56% (11.3) match
added_special_sparse modalities 3.3 23.5 ×7.04 ×1.01 8% (3.2) 0% (0.0) 11% (4.4) 82% (33.4) match
agentic-traces modalities 2.5 16.1 ×6.54 ×1.01 3% (1.8) 0% (0.0) 6% (3.5) 91% (55.7) match
agentic_swe modalities 2.5 24.4 ×9.81 ×1.00 3% (1.3) 0% (0.0) 6% (2.4) 91% (36.4) match
code_mixed modalities 2.5 20.4 ×8.21 ×1.01 3% (1.5) 0% (0.0) 6% (2.9) 91% (43.6) match
math_latex modalities 2.7 15.1 ×5.68 ×1.01 3% (1.9) 0% (0.0) 5% (3.4) 92% (59.7) match

Pre-tokenize: classify + fsm vs regex engines — ns/byte, lower better. The fsm is the scalar jump-table in both pipe columns; SIMD / scalar is the classify pass (regex pre-tokenizers have no SIMD fsm). ×vs = engine ÷ our pipeline (SIMD / scalar classify); onig & pcre2 (JIT) are C, fancy is pure-Rust fancy-regex, logos is a compile-time DFA lexer (approximate grammar; n/a for deepseek).

Fixture classify SIMD classify scalar pipe (SIMD cls + fsm) pipe (scalar cls + fsm) onig fancy pcre2 logos ×vs onig ×vs fancy ×vs pcre2 ×vs logos
amh_Ethi 2.78 3.44 3.68 4.34 28.2 21.6 5.8 4.8 7.7× / 6.5× 5.9× / 5.0× 1.6× / 1.3× 1.3× / 1.1×
arb_Arab 1.15 2.95 2.28 4.08 33.1 25.6 6.8 5.1 14.5× / 8.1× 11.2× / 6.3× 3.0× / 1.7× 2.2× / 1.2×
ben_Beng 1.61 2.90 3.02 4.31 66.1 56.9 13.5 4.1 21.9× / 15.4× 18.8× / 13.2× 4.5× / 3.1× 1.3× / 0.9×
cmn_Hani 1.12 2.24 2.31 3.42 25.9 22.1 5.8 2.3 11.2× / 7.6× 9.6× / 6.4× 2.5× / 1.7× 1.0× / 0.7×
ell_Grek 0.58 2.99 2.21 4.61 27.5 22.1 5.9 4.7 12.5× / 6.0× 10.0× / 4.8× 2.7× / 1.3× 2.1× / 1.0×
eng_Latn 0.10 1.45 3.09 4.44 45.0 41.4 11.9 3.7 14.6× / 10.1× 13.4× / 9.3× 3.8× / 2.7× 1.2× / 0.8×
heb_Hebr 1.15 3.03 2.30 4.19 28.6 26.8 6.8 3.0 12.4× / 6.8× 11.6× / 6.4× 3.0× / 1.6× 1.3× / 0.7×
hin_Deva 1.52 3.18 3.10 4.76 61.1 55.9 13.2 4.2 19.7× / 12.8× 18.0× / 11.7× 4.3× / 2.8× 1.4× / 0.9×
jpn_Jpan 1.69 3.33 2.11 3.75 23.6 17.4 5.0 3.8 11.2× / 6.3× 8.2× / 4.6× 2.3× / 1.3× 1.8× / 1.0×
kat_Geor 1.54 2.53 2.08 3.07 18.5 14.6 4.1 2.1 8.9× / 6.0× 7.0× / 4.7× 2.0× / 1.3× 1.0× / 0.7×
kor_Hang 1.22 2.72 2.48 3.97 31.9 27.4 7.5 3.6 12.9× / 8.0× 11.0× / 6.9× 3.0× / 1.9× 1.5× / 0.9×
rus_Cyrl 1.16 2.89 2.12 3.85 26.5 21.6 5.7 2.5 12.5× / 6.9× 10.2× / 5.6× 2.7× / 1.5× 1.2× / 0.7×
tam_Taml 0.93 2.96 2.64 4.67 71.0 60.5 14.1 3.8 26.9× / 15.2× 22.9× / 12.9× 5.3× / 3.0× 1.4× / 0.8×
tha_Thai 1.52 2.59 2.50 3.57 40.6 31.6 8.7 3.2 16.2× / 11.4× 12.6× / 8.8× 3.5× / 2.4× 1.3× / 0.9×
added_normalized_dense 0.06 1.48 1.01 2.42 24.0 23.3 6.4 1.8 23.8× / 9.9× 23.2× / 9.6× 6.4× / 2.6× 1.8× / 0.8×
added_normalized_sparse 0.06 1.77 2.00 3.71 31.8 31.5 8.6 2.6 15.9× / 8.6× 15.8× / 8.5× 4.3× / 2.3× 1.3× / 0.7×
added_special_dense 0.06 1.48 3.83 5.25 96.6 99.8 19.7 2.9 25.2× / 18.4× 26.1× / 19.0× 5.1× / 3.7× 0.8× / 0.6×
added_special_sparse 0.06 1.48 4.40 5.82 62.4 62.6 14.7 3.4 14.2× / 10.7× 14.2× / 10.8× 3.3× / 2.5× 0.8× / 0.6×
agentic-traces 0.71 1.48 3.52 4.28 59.6 62.6 15.1 4.6 16.9× / 13.9× 17.8× / 14.6× 4.3× / 3.5× 1.3× / 1.1×
agentic_swe 0.67 1.47 2.41 3.20 58.8 70.3 14.7 3.5 24.4× / 18.3× 29.2× / 22.0× 6.1× / 4.6× 1.4× / 1.1×
code_mixed 0.07 1.47 2.90 4.30 57.3 67.3 15.1 4.0 19.7× / 13.3× 23.2× / 15.6× 5.2× / 3.5× 1.4× / 0.9×
math_latex 0.71 1.48 3.38 4.15 51.3 51.8 13.8 4.1 15.2× / 12.4× 15.3× / 12.5× 4.1× / 3.3× 1.2× / 1.0×
gpt-oss — o200k-regex byte-level BPE (gpt-oss) · ×7.85 vs v0.23.1 · ×0.99 vs base gpt-oss speedup gpt-oss stage decomposition gpt-oss thread scaling

Memory (RSS MB, load+encode): v0.23.1 2+4 (peak 6) · Pipeline 3+0 (peak 6)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 4.1 62.7 ×15.23 ×1.01 4% (0.6) 0% (0.0) 30% (4.7) 65% (10.1) match
arb_Arab lang 4.1 24.3 ×6.01 ×0.95 1% (0.6) 0% (0.0) 8% (3.3) 91% (37.1) match
ben_Beng lang 4.8 28.2 ×5.82 ×0.98 2% (0.6) 0% (0.0) 11% (3.6) 88% (30.3) match
cmn_Hani lang 4.5 38.5 ×8.64 ×1.00 2% (0.6) 0% (0.0) 13% (3.2) 85% (21.5) match
ell_Grek lang 4.0 31.8 ×7.88 ×0.99 2% (0.6) 0% (0.0) 10% (3.2) 88% (27.1) match
eng_Latn lang 3.2 21.4 ×6.68 ×0.99 4% (2.0) 0% (0.0) 10% (4.4) 86% (39.5) match
heb_Hebr lang 3.8 27.7 ×7.25 ×0.98 2% (0.6) 0% (0.0) 9% (3.3) 89% (31.6) match
hin_Deva lang 4.6 23.5 ×5.14 ×0.97 1% (0.6) 0% (0.0) 9% (3.7) 90% (37.7) match
jpn_Jpan lang 4.8 37.3 ×7.76 ×1.00 2% (0.6) 0% (0.0) 12% (3.0) 86% (22.4) match
kat_Geor lang 5.1 26.7 ×5.28 ×1.01 2% (0.6) 0% (0.0) 8% (2.8) 91% (33.0) match
kor_Hang lang 3.5 46.0 ×13.07 ×0.98 3% (0.6) 0% (0.0) 17% (3.6) 80% (16.8) match
rus_Cyrl lang 4.3 22.6 ×5.32 ×0.98 1% (0.6) 0% (0.0) 7% (3.1) 92% (39.9) match
tam_Taml lang 5.2 30.7 ×5.90 ×1.00 2% (0.6) 0% (0.0) 10% (3.2) 88% (28.2) match
tha_Thai lang 5.9 28.2 ×4.79 ×1.02 2% (0.6) 0% (0.0) 9% (3.2) 89% (31.0) match
added_normalized_dense modalities 4.1 45.4 ×11.19 ×1.00 4% (0.8) 0% (0.0) 10% (2.1) 86% (17.6) match
added_normalized_sparse modalities 3.8 36.8 ×9.69 ×1.01 5% (1.3) 0% (0.0) 12% (3.1) 83% (21.5) match
added_special_dense modalities 3.6 55.0 ×15.15 ×1.01 29% (4.8) 1% (0.1) 29% (4.9) 42% (7.0) match
added_special_sparse modalities 3.6 34.6 ×9.67 ×1.01 12% (3.3) 0% (0.0) 21% (5.8) 68% (18.7) match
agentic-traces modalities 2.9 22.9 ×7.84 ×0.99 4% (1.8) 0% (0.0) 12% (5.0) 84% (36.3) match
agentic_swe modalities 3.0 22.1 ×7.35 ×1.00 3% (1.3) 0% (0.0) 8% (3.7) 89% (39.7) match
code_mixed modalities 3.1 30.0 ×9.56 ×0.99 5% (1.6) 0% (0.0) 13% (4.3) 82% (26.9) match
math_latex modalities 3.0 22.9 ×7.65 ×1.00 4% (1.9) 0% (0.0) 11% (4.7) 85% (36.4) match

Pre-tokenize: classify + fsm vs regex engines — ns/byte, lower better. The fsm is the scalar jump-table in both pipe columns; SIMD / scalar is the classify pass (regex pre-tokenizers have no SIMD fsm). ×vs = engine ÷ our pipeline (SIMD / scalar classify); onig & pcre2 (JIT) are C, fancy is pure-Rust fancy-regex, logos is a compile-time DFA lexer (approximate grammar; n/a for deepseek).

Fixture classify SIMD classify scalar pipe (SIMD cls + fsm) pipe (scalar cls + fsm) onig fancy pcre2 logos ×vs onig ×vs fancy ×vs pcre2 ×vs logos
amh_Ethi 2.77 3.43 4.72 5.38 29.4 13.8 7.0 4.6 6.2× / 5.5× 2.9× / 2.6× 1.5× / 1.3× 1.0× / 0.9×
arb_Arab 1.14 2.97 3.28 5.10 38.9 15.6 7.5 5.1 11.9× / 7.6× 4.8× / 3.1× 2.3× / 1.5× 1.6× / 1.0×
ben_Beng 1.62 2.90 3.63 4.92 23.8 10.6 5.4 2.7 6.5× / 4.8× 2.9× / 2.2× 1.5× / 1.1× 0.8× / 0.6×
cmn_Hani 1.12 2.24 3.25 4.37 22.3 10.7 5.5 2.4 6.9× / 5.1× 3.3× / 2.4× 1.7× / 1.3× 0.7× / 0.6×
ell_Grek 0.58 3.00 3.18 5.60 30.4 14.9 6.9 5.1 9.6× / 5.4× 4.7× / 2.7× 2.2× / 1.2× 1.6× / 0.9×
eng_Latn 0.11 1.46 4.44 5.79 44.4 28.3 13.6 4.0 10.0× / 7.7× 6.4× / 4.9× 3.1× / 2.4× 0.9× / 0.7×
heb_Hebr 1.14 3.01 3.29 5.16 33.2 16.8 8.1 2.8 10.1× / 6.4× 5.1× / 3.3× 2.4× / 1.6× 0.8× / 0.5×
hin_Deva 1.53 3.05 3.73 5.26 25.5 12.6 6.3 3.2 6.8× / 4.9× 3.4× / 2.4× 1.7× / 1.2× 0.9× / 0.6×
jpn_Jpan 1.69 3.41 3.03 4.75 20.4 9.2 4.8 3.7 6.7× / 4.3× 3.0× / 1.9× 1.6× / 1.0× 1.2× / 0.8×
kat_Geor 1.53 2.52 2.84 3.83 18.8 10.0 4.6 2.1 6.6× / 4.9× 3.5× / 2.6× 1.6× / 1.2× 0.7× / 0.5×
kor_Hang 1.22 2.68 3.63 5.08 34.0 18.7 8.9 3.6 9.4× / 6.7× 5.2× / 3.7× 2.5× / 1.8× 1.0× / 0.7×
rus_Cyrl 1.16 2.93 3.10 4.86 28.9 14.4 6.7 4.8 9.3× / 6.0× 4.7× / 3.0× 2.2× / 1.4× 1.6× / 1.0×
tam_Taml 0.93 2.94 3.16 5.18 19.0 8.6 4.2 3.0 6.0× / 3.7× 2.7× / 1.7× 1.3× / 0.8× 0.9× / 0.6×
tha_Thai 1.52 2.52 3.18 4.18 12.6 5.6 2.8 2.3 4.0× / 3.0× 1.8× / 1.3× 0.9× / 0.7× 0.7× / 0.6×
added_normalized_dense 0.06 1.77 2.13 3.83 34.4 17.0 11.3 2.0 16.2× / 9.0× 8.0× / 4.4× 5.3× / 2.9× 1.0× / 0.5×
added_normalized_sparse 0.06 1.48 3.09 4.51 36.6 20.0 11.7 2.8 11.8× / 8.1× 6.5× / 4.4× 3.8× / 2.6× 0.9× / 0.6×
added_special_dense 0.06 1.48 4.86 6.28 87.4 66.8 24.0 3.2 18.0× / 13.9× 13.7× / 10.6× 4.9× / 3.8× 0.7× / 0.5×
added_special_sparse 0.06 1.48 5.83 7.25 58.5 41.2 16.6 3.4 10.0× / 8.1× 7.1× / 5.7× 2.9× / 2.3× 0.6× / 0.5×
agentic-traces 0.72 1.50 4.96 5.75 53.1 40.8 16.7 4.8 10.7× / 9.2× 8.2× / 7.1× 3.4× / 2.9× 1.0× / 0.8×
agentic_swe 0.66 1.44 3.71 4.50 54.8 48.3 17.1 3.5 14.8× / 12.2× 13.0× / 10.8× 4.6× / 3.8× 1.0× / 0.8×
code_mixed 0.09 1.44 4.33 5.67 53.5 46.4 17.2 4.2 12.3× / 9.4× 10.7× / 8.2× 4.0× / 3.0× 1.0× / 0.7×
math_latex 0.75 1.48 4.73 5.46 52.8 34.8 15.7 4.4 11.2× / 9.7× 7.3× / 6.4× 3.3× / 2.9× 0.9× / 0.8×
glm-5.2 — cl100k-variant regex byte-level BPE (glm-5.2) · ×13.12 vs v0.23.1 · ×0.99 vs base glm-5.2 speedup glm-5.2 stage decomposition glm-5.2 thread scaling

Memory (RSS MB, load+encode): v0.23.1 2+4 (peak 6) · Pipeline 3+0 (peak 6)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 4.6 60.8 ×13.37 ×0.94 8% (1.2) 0% (0.0) 23% (3.7) 69% (10.9) match
arb_Arab lang 3.8 73.8 ×19.21 ×0.97 9% (1.2) 0% (0.0) 18% (2.3) 73% (9.5) match
ben_Beng lang 3.2 68.3 ×21.57 ×0.97 8% (1.2) 0% (0.0) 22% (3.1) 70% (9.9) match
cmn_Hani lang 4.6 70.1 ×15.33 ×1.05 9% (1.2) 0% (0.0) 17% (2.4) 74% (10.1) match
ell_Grek lang 4.0 70.9 ×17.88 ×0.91 9% (1.2) 0% (0.0) 16% (2.2) 75% (10.1) match
eng_Latn lang 3.3 23.0 ×6.89 ×1.00 6% (2.5) 0% (0.0) 7% (2.9) 87% (38.0) match
heb_Hebr lang 3.8 69.4 ×18.05 ×0.92 8% (1.2) 0% (0.0) 17% (2.3) 75% (10.4) match
hin_Deva lang 3.0 64.0 ×21.27 ×0.94 8% (1.2) 0% (0.0) 21% (3.2) 71% (10.7) match
jpn_Jpan lang 4.6 70.4 ×15.15 ×1.01 9% (1.2) 0% (0.0) 16% (2.2) 75% (10.1) match
kat_Geor lang 4.5 84.7 ×18.64 ×0.95 10% (1.2) 0% (0.0) 18% (2.1) 71% (8.1) match
kor_Hang lang 3.7 70.3 ×19.07 ×1.04 9% (1.2) 0% (0.0) 18% (2.5) 73% (9.9) match
rus_Cyrl lang 4.1 40.8 ×10.05 ×1.04 5% (1.2) 0% (0.0) 9% (2.1) 86% (20.3) match
tam_Taml lang 3.1 68.0 ×21.60 ×0.93 8% (1.2) 0% (0.0) 19% (2.7) 73% (10.5) match
tha_Thai lang 3.9 71.8 ×18.50 ×1.02 9% (1.2) 0% (0.0) 19% (2.5) 72% (9.5) match
added_normalized_dense modalities 4.5 46.2 ×10.22 ×1.01 7% (1.3) 0% (0.0) 6% (1.1) 88% (17.9) match
added_normalized_sparse modalities 4.1 40.6 ×10.00 ×1.01 8% (1.9) 0% (0.0) 8% (1.9) 84% (19.8) match
added_special_dense modalities 3.8 41.5 ×11.05 ×1.02 52% (12.0) 0% (0.0) 20% (4.7) 27% (6.3) match
added_special_sparse modalities 3.7 34.3 ×9.28 ×1.01 24% (6.7) 1% (0.2) 16% (4.4) 60% (16.8) match
agentic-traces modalities 3.0 24.0 ×7.89 ×1.00 6% (2.4) 0% (0.0) 9% (3.7) 85% (35.0) match
agentic_swe modalities 3.1 22.4 ×7.27 ×1.00 4% (1.9) 0% (0.0) 7% (2.9) 89% (39.1) match
code_mixed modalities 3.3 31.0 ×9.39 ×0.98 7% (2.2) 0% (0.0) 10% (3.2) 83% (26.2) match
math_latex modalities 3.1 24.5 ×7.98 ×1.00 6% (2.5) 0% (0.0) 9% (3.4) 85% (34.2) match

Pre-tokenize: classify + fsm vs regex engines — ns/byte, lower better. The fsm is the scalar jump-table in both pipe columns; SIMD / scalar is the classify pass (regex pre-tokenizers have no SIMD fsm). ×vs = engine ÷ our pipeline (SIMD / scalar classify); onig & pcre2 (JIT) are C, fancy is pure-Rust fancy-regex, logos is a compile-time DFA lexer (approximate grammar; n/a for deepseek).

Fixture classify SIMD classify scalar pipe (SIMD cls + fsm) pipe (scalar cls + fsm) onig fancy pcre2 logos ×vs onig ×vs fancy ×vs pcre2 ×vs logos
amh_Ethi 2.77 3.40 3.68 4.31 27.3 15.8 5.8 4.8 7.4× / 6.3× 4.3× / 3.7× 1.6× / 1.3× 1.3× / 1.1×
arb_Arab 1.15 2.96 2.29 4.10 28.3 19.1 6.8 5.2 12.4× / 6.9× 8.3× / 4.7× 3.0× / 1.7× 2.3× / 1.3×
ben_Beng 1.61 2.95 3.07 4.41 45.9 27.3 10.1 3.7 15.0× / 10.4× 8.9× / 6.2× 3.3× / 2.3× 1.2× / 0.8×
cmn_Hani 1.12 2.26 2.35 3.49 19.7 11.6 4.6 2.4 8.4× / 5.6× 4.9× / 3.3× 2.0× / 1.3× 1.0× / 0.7×
ell_Grek 0.58 2.95 2.20 4.57 28.8 16.7 6.0 4.8 13.0× / 6.3× 7.6× / 3.7× 2.7× / 1.3× 2.2× / 1.0×
eng_Latn 0.09 1.46 2.95 4.31 46.1 31.9 12.2 3.7 15.6× / 10.7× 10.8× / 7.4× 4.1× / 2.8× 1.3× / 0.9×
heb_Hebr 1.15 3.02 2.30 4.17 31.6 19.3 6.9 3.1 13.7× / 7.6× 8.4× / 4.6× 3.0× / 1.6× 1.3× / 0.7×
hin_Deva 1.51 3.09 3.17 4.75 47.6 30.2 10.9 4.0 15.0× / 10.0× 9.5× / 6.4× 3.4× / 2.3× 1.3× / 0.9×
jpn_Jpan 1.69 3.40 2.16 3.87 18.8 10.0 4.0 3.7 8.7× / 4.9× 4.6× / 2.6× 1.9× / 1.0× 1.7× / 1.0×
kat_Geor 1.54 2.51 2.08 3.05 17.3 11.0 4.1 2.0 8.3× / 5.7× 5.3× / 3.6× 2.0× / 1.4× 1.0× / 0.7×
kor_Hang 1.23 2.70 2.49 3.96 31.7 21.1 7.5 3.6 12.7× / 8.0× 8.5× / 5.3× 3.0× / 1.9× 1.4× / 0.9×
rus_Cyrl 1.16 2.92 2.13 3.89 26.5 16.2 6.0 2.6 12.4× / 6.8× 7.6× / 4.2× 2.8× / 1.5× 1.2× / 0.7×
tam_Taml 0.93 2.91 2.65 4.63 44.8 26.0 9.8 3.4 16.9× / 9.7× 9.8× / 5.6× 3.7× / 2.1× 1.3× / 0.7×
tha_Thai 1.55 2.62 2.53 3.61 28.3 15.8 6.7 3.0 11.2× / 7.8× 6.2× / 4.4× 2.6× / 1.8× 1.2× / 0.8×
added_normalized_dense 0.06 1.47 1.12 2.54 24.0 16.9 6.6 1.8 21.4× / 9.4× 15.1× / 6.7× 5.9× / 2.6× 1.6× / 0.7×
added_normalized_sparse 0.06 1.47 1.88 3.30 31.5 22.5 9.0 2.5 16.7× / 9.5× 12.0× / 6.8× 4.8× / 2.7× 1.3× / 0.8×
added_special_dense 0.06 1.47 4.70 6.11 93.8 71.4 20.9 3.1 20.0× / 15.3× 15.2× / 11.7× 4.4× / 3.4× 0.7× / 0.5×
added_special_sparse 0.06 1.47 4.42 5.84 62.4 44.8 15.2 3.6 14.1× / 10.7× 10.1× / 7.7× 3.4× / 2.6× 0.8× / 0.6×
agentic-traces 0.72 1.47 3.66 4.41 51.8 42.8 14.8 4.5 14.2× / 11.8× 11.7× / 9.7× 4.0× / 3.3× 1.2× / 1.0×
agentic_swe 0.66 1.45 2.86 3.66 55.2 50.7 15.4 3.4 19.3× / 15.1× 17.7× / 13.9× 5.4× / 4.2× 1.2× / 0.9×
code_mixed 0.07 1.44 3.20 4.57 53.9 48.2 15.0 3.9 16.8× / 11.8× 15.1× / 10.5× 4.7× / 3.3× 1.2× / 0.9×
math_latex 0.70 1.47 3.41 4.18 52.5 39.3 14.0 4.1 15.4× / 12.5× 11.5× / 9.4× 4.1× / 3.3× 1.2× / 1.0×
llama-2 — model-bounded BPE, no pre-tokenizer · ×4.12 vs v0.23.1 · ×1.02 vs base llama-2 speedup llama-2 stage decomposition llama-2 thread scaling

Memory (RSS MB, load+encode): v0.23.1 15+0 (peak 15) · Pipeline 14+0 (peak 14)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 4.3 47.7 ×11.03 ×1.01 0% (0.0) 14% (2.6) 0% (0.0) 86% (16.8) match
arb_Arab lang 9.3 53.3 ×5.72 ×1.01 0% (0.0) 17% (3.1) 0% (0.0) 82% (14.8) match
ben_Beng lang 10.0 93.2 ×9.35 ×1.01 0% (0.0) 20% (2.0) 0% (0.0) 79% (8.1) match
cmn_Hani lang 8.5 72.9 ×8.61 ×1.01 0% (0.1) 3% (0.4) 0% (0.0) 97% (12.1) match
ell_Grek lang 9.3 61.3 ×6.61 ×0.99 0% (0.0) 20% (3.0) 0% (0.0) 80% (12.4) match
eng_Latn lang 4.3 6.6 ×1.54 ×1.04 0% (0.1) 5% (6.7) 0% (0.0) 95% (141.4) match
heb_Hebr lang 9.1 66.1 ×7.25 ×1.00 0% (0.0) 22% (3.2) 0% (0.0) 78% (11.3) match
hin_Deva lang 11.1 86.6 ×7.78 ×1.00 0% (0.0) 25% (2.7) 0% (0.0) 74% (8.2) match
jpn_Jpan lang 12.0 96.6 ×8.06 ×1.04 1% (0.1) 3% (0.3) 0% (0.0) 96% (9.3) match
kat_Geor lang 13.0 100.6 ×7.74 ×1.00 0% (0.0) 19% (1.8) 0% (0.0) 81% (7.7) match
kor_Hang lang 6.7 57.9 ×8.61 ×1.04 0% (0.1) 19% (3.2) 0% (0.0) 80% (13.3) match
rus_Cyrl lang 8.5 17.0 ×2.00 ×1.01 0% (0.0) 5% (2.7) 0% (0.1) 95% (55.2) match
tam_Taml lang 11.4 101.5 ×8.92 ×1.07 0% (0.0) 17% (1.6) 0% (0.0) 82% (7.7) match
tha_Thai lang 14.3 96.2 ×6.72 ×1.00 0% (0.0) 9% (0.9) 0% (0.0) 91% (9.0) match
added_normalized_dense modalities 5.3 8.9 ×1.66 ×0.99 0% (0.0) 3% (3.7) 0% (0.0) 97% (107.5) match
added_normalized_sparse modalities 4.8 7.6 ×1.59 ×1.05 0% (0.1) 5% (5.9) 0% (0.0) 95% (120.7) match
added_special_dense modalities 4.6 20.5 ×4.42 ×1.03 10% (5.0) 34% (16.0) 2% (1.0) 54% (25.8) match
added_special_sparse modalities 6.7 10.4 ×1.55 ×0.99 2% (2.1) 14% (13.5) 0% (0.4) 83% (77.8) match
agentic-traces modalities 4.4 7.5 ×1.70 ×1.05 0% (0.1) 5% (6.0) 0% (0.0) 95% (126.3) match
agentic_swe modalities 4.0 7.8 ×1.97 ×1.05 0% (0.0) 7% (9.5) 0% (0.0) 93% (118.3) match
code_mixed modalities 4.1 7.5 ×1.81 ×1.04 0% (0.0) 6% (8.2) 0% (0.0) 94% (124.1) match
math_latex modalities 4.3 7.1 ×1.65 ×1.05 0% (0.1) 4% (6.2) 0% (0.0) 95% (133.3) match
llama-3 — cl100k-regex byte-level BPE (llama-3), single regex · ×9.09 vs v0.23.1 · ×1.02 vs base llama-3 speedup llama-3 stage decomposition llama-3 thread scaling

Memory (RSS MB, load+encode): v0.23.1 129+0 (peak 189) · Pipeline 130+0 (peak 189)

Fixture Group v0.23.1 MB/s Pipeline MB/s Speedup Δ base added-token normalize pre-tokenize model Ids
amh_Ethi lang 3.2 52.1 ×16.15 ×1.01 4% (0.7) 0% (0.0) 21% (3.7) 75% (12.7) match
arb_Arab lang 3.8 18.3 ×4.82 ×1.01 1% (0.6) 0% (0.0) 4% (2.3) 95% (50.4) match
ben_Beng lang 2.9 32.0 ×11.06 ×0.96 2% (0.6) 0% (0.0) 10% (3.1) 88% (26.7) match
cmn_Hani lang 4.1 17.6 ×4.25 ×1.06 1% (0.6) 0% (0.0) 4% (2.3) 94% (50.5) match
ell_Grek lang 4.1 20.3 ×5.00 ×1.00 1% (0.6) 0% (0.0) 5% (2.2) 94% (44.8) match
eng_Latn lang 3.5 44.3 ×12.56 ×1.12 10% (2.0) 0% (0.0) 16% (2.9) 74% (13.7) match
heb_Hebr lang 3.3 27.2 ×8.33 ×1.00 2% (0.6) 0% (0.0) 6% (2.3) 92% (33.1) match
hin_Deva lang 3.7 79.3 ×21.31 ×1.01 5% (0.6) 0% (0.0) 28% (3.2) 67% (7.7) match
jpn_Jpan lang 4.4 17.2 ×3.89 ×1.03 1% (0.6) 0% (0.0) 4% (2.1) 95% (52.6) match
kat_Geor lang 3.9 37.4 ×9.67 ×1.00 2% (0.6) 0% (0.0) 8% (2.1) 89% (22.6) match
kor_Hang lang 3.5 19.9 ×5.62 ×1.03 1% (0.6) 0% (0.0) 5% (2.5) 94% (44.5) match
rus_Cyrl lang 4.1 17.1 ×4.22 ×1.05 1% (0.6) 0% (0.0) 4% (2.1) 95% (53.1) match
tam_Taml lang 2.9 36.3 ×12.55 ×0.98 2% (0.6) 0% (0.0) 10% (2.7) 88% (23.4) match
tha_Thai lang 4.4 21.5 ×4.89 ×1.01 1% (0.6) 0% (0.0) 6% (2.5) 93% (41.6) match
added_normalized_dense modalities 3.9 27.7 ×7.05 ×1.03 2% (0.8) 0% (0.0) 3% (1.0) 95% (32.9) match
added_normalized_sparse modalities 4.1 43.6 ×10.61 ×1.03 6% (1.3) 0% (0.0) 9% (1.9) 86% (18.6) match
added_special_dense modalities 3.4 72.3 ×21.19 ×0.99 41% (5.0) 0% (0.0) 40% (4.9) 20% (2.5) match
added_special_sparse modalities 3.7 70.6 ×19.05 ×1.01 25% (3.2) 0% (0.0) 38% (4.8) 37% (4.7) match
agentic-traces modalities 3.2 36.9 ×11.60 ×1.02 7% (1.7) 0% (0.0) 15% (3.6) 78% (18.7) match
agentic_swe modalities 3.2 28.7 ×8.95 ×1.07 4% (1.3) 0% (0.0) 9% (2.8) 87% (28.3) match
code_mixed modalities 3.4 47.1 ×13.92 ×1.05 8% (1.5) 0% (0.0) 17% (3.2) 75% (13.9) match
math_latex modalities 3.2 40.2 ×12.63 ×1.06 9% (1.9) 0% (0.0) 16% (3.4) 75% (16.1) match

Pre-tokenize: classify + fsm vs regex engines — ns/byte, lower better. The fsm is the scalar jump-table in both pipe columns; SIMD / scalar is the classify pass (regex pre-tokenizers have no SIMD fsm). ×vs = engine ÷ our pipeline (SIMD / scalar classify); onig & pcre2 (JIT) are C, fancy is pure-Rust fancy-regex, logos is a compile-time DFA lexer (approximate grammar; n/a for deepseek).

Fixture classify SIMD classify scalar pipe (SIMD cls + fsm) pipe (scalar cls + fsm) onig fancy pcre2 logos ×vs onig ×vs fancy ×vs pcre2 ×vs logos
amh_Ethi 2.84 3.42 3.66 4.24 27.9 15.9 5.9 4.8 7.6× / 6.6× 4.3× / 3.8× 1.6× / 1.4× 1.3× / 1.1×
arb_Arab 1.15 2.98 2.27 4.10 31.6 18.9 6.8 5.2 13.9× / 7.7× 8.3× / 4.6× 3.0× / 1.7× 2.3× / 1.3×
ben_Beng 1.61 2.95 3.06 4.39 40.2 26.9 10.1 3.7 13.2× / 9.2× 8.8× / 6.1× 3.3× / 2.3× 1.2× / 0.8×
cmn_Hani 1.12 2.25 2.34 3.47 19.5 11.4 4.7 2.3 8.4× / 5.6× 4.9× / 3.3× 2.0× / 1.3× 1.0× / 0.7×
ell_Grek 0.58 2.94 2.18 4.54 28.7 16.8 6.2 4.7 13.2× / 6.3× 7.7× / 3.7× 2.9× / 1.4× 2.2× / 1.0×
eng_Latn 0.09 1.46 2.91 4.28 45.2 31.9 12.4 3.7 15.5× / 10.6× 11.0× / 7.5× 4.2× / 2.9× 1.3× / 0.9×
heb_Hebr 1.14 3.04 2.28 4.18 31.1 19.7 6.9 3.0 13.6× / 7.4× 8.6× / 4.7× 3.0× / 1.6× 1.3× / 0.7×
hin_Deva 1.51 3.07 3.18 4.74 49.1 30.5 11.3 4.0 15.4× / 10.3× 9.6× / 6.4× 3.6× / 2.4× 1.3× / 0.9×
jpn_Jpan 1.69 3.31 2.14 3.77 18.6 9.6 4.0 3.8 8.7× / 4.9× 4.5× / 2.5× 1.9× / 1.1× 1.8× / 1.0×
kat_Geor 1.54 2.54 2.07 3.07 16.7 11.1 4.2 2.0 8.1× / 5.4× 5.3× / 3.6× 2.0× / 1.4× 1.0× / 0.6×
kor_Hang 1.23 2.73 2.47 3.97 30.4 20.6 7.4 3.6 12.3× / 7.7× 8.3× / 5.2× 3.0× / 1.9× 1.5× / 0.9×
rus_Cyrl 1.16 2.89 2.12 3.84 26.6 16.3 5.9 2.5 12.5× / 6.9× 7.7× / 4.2× 2.8× / 1.5× 1.2× / 0.7×
tam_Taml 0.93 2.94 2.66 4.67 43.6 26.6 9.8 3.5 16.4× / 9.3× 10.0× / 5.7× 3.7× / 2.1× 1.3× / 0.8×
tha_Thai 1.51 2.54 2.52 3.55 28.0 15.9 6.7 3.0 11.1× / 7.9× 6.3× / 4.5× 2.7× / 1.9× 1.2× / 0.8×
added_normalized_dense 0.06 1.47 1.03 2.45 23.3 16.7 6.6 1.8 22.7× / 9.5× 16.3× / 6.8× 6.4× / 2.7× 1.8× / 0.8×
added_normalized_sparse 0.06 1.77 1.86 3.57 30.7 22.6 9.0 2.5 16.5× / 8.6× 12.2× / 6.3× 4.9× / 2.5× 1.4× / 0.7×
added_special_dense 0.06 1.47 4.87 6.28 92.4 72.8 20.7 3.1 19.0× / 14.7× 15.0× / 11.6× 4.3× / 3.3× 0.6× / 0.5×
added_special_sparse 0.06 1.47 4.80 6.21 60.0 44.5 15.1 3.5 12.5× / 9.7× 9.3× / 7.2× 3.2× / 2.4× 0.7× / 0.6×
agentic-traces 0.74 1.47 3.65 4.39 52.0 42.7 14.9 4.5 14.2× / 11.9× 11.7× / 9.7× 4.1× / 3.4× 1.2× / 1.0×
agentic_swe 0.66 1.44 2.83 3.61 54.4 49.2 15.4 3.4 19.2× / 15.1× 17.4× / 13.6× 5.5× / 4.3× 1.2× / 0.9×
code_mixed 0.10 1.44 3.17 4.51 52.1 49.4 15.2 4.0 16.5× / 11.6× 15.6× / 10.9× 4.8× / 3.4× 1.3× / 0.9×
math_latex 0.71 1.48 3.40 4.17 50.4 37.3 14.0 4.1 14.8× / 12.1× 11.0× / 8.9× 4.1× / 3.4× 1.2× / 1.0×
Not yet supported: t5-base
t5-base not supported

ArthurZucker and others added 16 commits July 15, 2026 19:25
Ports the self-contained pure-Rust normalization port from feat/fast-normalization:
nfd.rs + nfd_tables/nfkd_tables/compose_tables, wired additively into lib.rs. No
change to atomsplit's classify/fsm/simd core. Differential tests vs unicode-normalization
(dev-only) pass byte-exact.
Rewire the unicode normalizers off the unicode-normalization crate onto the pure-Rust
atomsplit::nfd port. Byte-exact (normalizers::unicode tests pass).
Make classify/classify_scalar + the NEON kernel generic over <const CONT, MB,
CJK_TAG> + a &Tables arg (no trait — a different Tables instance is a different
scheme; CJK_TAG=NO_CJK compiles out the CJK lead-byte shortcut). Pretok Atom path
is a thin behaviour-identical wrapper (parity/fsm stay byte-exact). Add the norm
scheme: NORM_TABLES + norm_classify::classify, with a SIMD==scalar parity test.
Thread <const CONT, MB, CJK_TAG> + &Tables through the AVX-512/SSE x86_body! macro
(via macro args, C/M/CJ generics to avoid the const-in-pattern collision) and the
wasm SIMD128 kernel, matching the NEON generic-ization. Cross-checked:
x86_64-unknown-linux-gnu + wasm32-unknown-unknown build clean; aarch64 unchanged.
…split::norm_classify

Consume the landed norm_classify SIMD engine: one classify pass tags every byte
with the NormClass bitmask, then copy INERT runs verbatim and touch only flagged
chars (clean_text/CJK/lowercase/strip_accents). New tagged.rs helper (tag_driven);
Bert fast path + Strip/StripAccents + Lowercase rewired onto it. Added a scalar
norm_classify::classify_scalar wrapper for the differential test.

Byte-exact vs the legacy normalizers (tag_driven_bert_matches_legacy over all
flag combos + strip/lowercase differential tests). Out-of-scope Sequence
metaspace-fusion from the PR was NOT ported (needs fancy-regex + diverged).
…+ fix NFD 2-byte skip panic

Adds `benches/normalize.rs`: NFD/NFKD/NFC/NFKC throughput for atomsplit's `nfd` against
`unicode-normalization` (Rust reference / test oracle) and, behind `--features xxutf`, the
dzfrias/xxUTF C+SIMD normalizer (a feature-gated build script downloads + compiles the pinned MIT
amalgamation; default builds stay C-free — no build script, no `cc`, no network). Same corpora as the
regex bench; min-of-7 ns/byte + a byte-exactness gate vs unicode-normalization.

The bench immediately surfaced a panic: `skip_clear_2byte` (the vld2q stable-run skip) advanced `i` by
64 per fully-clear 32-byte window — a stray `i += 32` after the `match` on top of the `None => i += 32`
arm — while verifying only 32, desyncing the byte cursor from char boundaries when the unverified half
held a width-≠2 char, so `&input[i..ns]` in `decompose` sliced mid-codepoint and panicked (Cyrillic/Greek
NFD/NFKD). The 3-byte twin `skip_clear_3byte` is correct (single `None => i += 48`, no trailing add);
dropped the extra increment to match. nfd differential tests still pass (byte-exact vs unicode-normalization).
…ose only mark-clusters

`compose` recomposed the whole string once ANY composition-relevant char appeared. Instead, stream: copy
every STABLE run verbatim (SIMD-skipped by `next_set::<R>`) and decompose+recompose only the ACTIVE
window around each relevant char — `[starter-before-rel, next-not-in-R-char)`. A char not in R is ccc-0
& QC=Yes: it neither composes backward nor is reached by a preceding composition, so it's a safe segment
boundary. Byte-exact (nfd differential tests + the normalize bench's exactness gate on 11 corpora × 4 forms).

Turns NFC/NFKC on marked scripts from whole-document recompose into ~80%-copy: NFC Russian 7.9→0.58 ns/B,
Hebrew 7.7→0.75, Arabic 8.3→0.85, Korean 3.8→0.91, Greek 11.6→0.58. vs xxUTF (C+SIMD), NFC now ties on
Russian/Greek, wins on Chinese/Japanese/French, and is within ~1.2× on Hebrew/Arabic (dense Devanagari/
Thai marks remain ~0.4× — those and all NFD/NFKD need a SIMD decompose kernel). No change to NFD/NFKD.
…atch or beat xxUTF

Three layers, all byte-exact vs unicode-normalization (exhaustive unit gates: every cp × mark
suffixes × 4 forms through BOTH kernel paths, plus 11-corpus bench exactness):

1. Byte-form decomposition tables (generated, derived 1:1 from the committed (ccc,char) tries —
   `gen_byte_tables`): each unstable cp's UTF-8 decomposition blob + [first_ccc, last_ccc,
   mark_run_off] headers. The owned path emits via ONE unaligned 16-byte copy per char (arithmetic
   direct-store jamo for Hangul) instead of char-encode pushes; NFKC additionally bakes composed
   blobs for neighbour-inert compat starters (fullwidth punctuation → one copy, no window).

2. A dedicated check-tag classify scheme (`check_tag`, generated into NFD_CHECK_TABLES via
   bitmap_gen, which now derives it from atomsplit's own tables): per-cp tag = ccc RANK
   (order-preserving, 6 bits) | QC-Maybe flag (0x40) | compat-change (0x3C/0x3D) | canonical-change
   split by composition-relevance (0x7E stable / 0x7D relevant). Both decompose checks and the
   NFC/NFKC quick-check run on pure tag arithmetic — no decode, no table gathers; composed
   é/ά/Hangul are excluded from the compose scan entirely.

3. Strategy dispatch by a 64-byte content sample: the tag kernel for marked scripts (flat
   whole-buffer scan, candidates verified in-register), the bitset kernel for ASCII-dominant and
   CJK-block text — with fused ASCII scan+copy, a byte-compare Hangul-or-ASCII run skip (Korean
   NFC 13.4 → 0.19 ns/B, 3× FASTER than xxUTF), verify-as-you-go owned rebuilds (in-order marks
   stay in the gap memcpy), and callless small-copies.

vs xxUTF (C+SIMD, "fastest open source"), 44 cells (4 forms × 11 scripts): 38 tie-or-beat
(NFC/NFKC sweep almost everything at 1.2-3×; English 1.7-1.9×, Russian/Greek/Hebrew/Arabic/Hindi
1.0-1.4×, CJK NFC 1.4-3×). Remaining 0.7-0.95×: Thai/French/Korean NFD-side + NFKC Chinese —
bounded by the classify pass on those scripts. vs unicode-normalization: 6-500×.
…match or beat xxUTF

Byte-exact throughout (exhaustive per-cp gates through every kernel path + 11-corpus exactness),
0 warnings. Final bench vs xxUTF (C+SIMD): min 1.0×, max 3.6×, mean 1.4×; vs unicode-normalization
4-500×. The last mile:

- 0x7D/0x7E tag split (composition-relevant vs composition-stable canonical-decomposers): composed
  é/ά/Hangul become tag-only compose skips; scan excludes 0x7E entirely (Korean NFC 13.4 → 0.19
  ns/B — 3× faster than xxUTF; Greek/French NFC to 1.2-1.3×).
- Baked NFKC blobs for neighbour-inert compat starters, emitted in both compose paths (NFKC Chinese
  1.82 → 0.44, fullwidth punctuation is one 16-byte copy).
- Hangul-or-ASCII byte-class run skip (16-byte chunks, pair-constrained lead ranges — no decode).
- Streaming in-register check for single-block 3-byte scripts (Thai/Lao): the block's own fast3
  check-tag table probed via vqtbl4 with vectorized break + ccc-rank-order tests, block-cached
  across ASCII holes, plus its verify-as-you-go owned twin (Thai NFD 0.96 → 0.47, tie) and a
  compose-predicate variant (Thai NFC 0.61 → 0.41, 1.5×). A scalar referee using the identical tag
  semantics keeps every bail byte-exact.
- Write-through skips: the fused ASCII loop stores its test register and movemask-jumps to the
  first non-ASCII byte (French NFD/NFKD → 1.0-1.1×); the 3-byte skip vst3q's its own deinterleaved
  registers back and forwards the decoded SET codepoint it stops at (Japanese NFD/NFKD 0.63 → 0.45,
  1.3×; lifted Chinese NFKD and Korean too).
- Hangul syllable-run subloop in the emit (direct 3-byte decode, no generic continuation glue) —
  Korean NFD/NFKD to 1.1-1.2×.
…xUTF

Thread a `const SIMD: bool` through the drivers (dead-code eliminated — zero cost on the shipping
path) and expose `#[doc(hidden)] atomnorm::scalar::{nfd,nfkd,nfc,nfkc}`. The bench now shows all
four columns and the exactness gate covers BOTH paths; the exhaustive tests differentially verify
the scalar path too.

Notable: the scalar path alone runs 0.27-1.6 ns/B (3-20x over unicode-normalization) and even beats
the SIMD path on hit-dense scripts (Greek/Hindi) — the layered-skip design carries most of the win,
SIMD is the cherry on top for clean-run scripts (English/French/Russian 5-10x further).
…in atomnorm

Restore atomsplit/ and bitmap_gen/ byte-for-byte to the PR base
(feat/train_encode_split): drop norm_classify.rs, norm_tables.rs, and the
classify_with genericization of the classify engine. Revert the tk-encode
tagged-normalizer fast paths (bert/strip/utils/tagged) that depended on them;
that work stays on feat/fast-normalization. tk-encode keeps only the atomnorm
integration (unicode.rs NFD/NFKD/NFC/NFKC).
…, test the NEON path on an arm runner instead of SDE/wasm jobs atomnorm has no code for
…fused bert (2-134x)

One new primitive covers all four: a layered skip-scan over per-rule property
bitmaps (runtime-OR'd once per config, lead mask derived), with an in-register
ASCII lane — the PR #2036 |0x20 port plus bert/nmt whitespace folds — and
scalar-first streak escalation into the NEON kernel so hit-dense scripts never
pay kernel round-trips. Property sets are generated from the SAME crates the
legacy tk-encode normalizers use (std to_lowercase, unicode_categories,
unicode-normalization-alignments), so they are bug-compatible by construction;
bert's strip fixup routes unstable clusters through the existing NFD machinery
(cross-char reorder + clean-removed transparency stay exact).

tk-encode pipeline impls repointed to atomnorm; legacy NormalizedString paths
untouched. Exhaustive per-codepoint tests (all 1.1M cps in context, both
paths) + all 16 bert flag combos; 297 tk-encode tests green.

Wikipedia bench (ns/B, simd vs legacy char-iterator): bert english 15.9->0.12
(134x), chinese 17.9->2.5 (7.3x), all cells 2.1-134x; nmt 0.11-0.17 on
alphabetic scripts (12-17x); lowercase 1.3-48x; strip_accents 0.5-39x
(mark-dense scripts stay ~legacy speed, <=2 ns/B absolute — known ceiling).
@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

…8-35x, byte-exact)

atomnorm grows a public runtime Scanner — the parameterized skip-scan over a
caller-provided set (with an ASCII-membership lane the built-ins don't need).
tk-encode wraps spm_precompiled (serde-transparent) and derives at model load:
scan-hot = single-char keys ∪ multi-char-key tails (own DoubleArray DFS over
the charsmap blob), cluster-class = probed from the same unicode_segmentation
that runs the exact walk. Scan skips cold runs; a hit walks back over class
chars to a provable grapheme boundary and runs the exact walk locally. The
reachability invariant (every multi-char key tail is cluster-class) is
verified per charsmap — violation falls back to the plain walk.

albert nmt_nfkc on Wikipedia corpora: english 16.6->0.64 ns/B (26x), russian
12.3->0.35 (35x), korean 10.7->1.01 (10.6x), chinese 8.2->2.9 (2.8x, genuine
fullwidth remapping). spm_precompiled stays the exact engine.
…-35x, bert 5-25x on cased scripts)

Source table = SCAN_REG2 (88 cps: 2-byte uppercase mapping to exactly cp+0x20 —
Latin-1/Greek/Cyrillic capitals), masked per config against every other enabled
rule (under bert strip, decomposing capitals like Й stay scalar fixups). Target
is arithmetic: continuation +0x20 with the UTF-8 carry (-0x40, lead +1) so
Р D0A0 -> р D180 in-register. New write-only scan2_case NEON kernel (skip2-style
dual 2048-bit probes: SET stops, REG2 transforms; trailing leads deferred so
pairs never straddle chunks) + the scalar twin; next_hit escalates into it when
suspect-2-byte probes dominate, clean terrain keeps the light 32B kernel.

Also fixes lead-mask over-coverage: lead E0 valid cps start at 0x800, so Latin-1
capitals no longer taint the whole Devanagari/Thai range (lower hindi 2.0->0.07).

lower ns/B vs legacy: russian 0.43 (8.8x, was 1.9x), greek 0.41 (9.0x), hindi
0.071 (31x), thai 0.065 (29x); bert: russian 0.89 (14.1x, was 5.5x), hebrew
0.49 (24.8x), arabic 1.22 (9.7x). All exhaustive per-cp suites green, both paths.
…er (forms, lower, strip, nmt, bert, spm-precompiled) vs its legacy implementation on the Wikipedia corpora
…ed walk

simd_norm: delete the dead vld2q experiment in skip3 (if false block), drop the
per-body cfg(aarch64) (the module is already gated in lib.rs), and factor the
shared pieces — first_lane (POW movemask), class2, load2048/probe2048 (2048-bit
in-register probe), upper_mask, ascii_policy (the bert/nmt fold+removal masks,
previously duplicated verbatim in scan_prefix and scan2_case). skip2_ascii
stays deliberately hand-inlined: the helper-factored variant (array-of-tables)
measured +25% on NFD Russian — with one table the four vqtbl4 groups must stay
individual SSA values; noted in a comment.

precompiled: factor transform_grapheme — the one copy of the exact-walk step
both normalize_scan and normalize_walk share (was duplicated with subtle
prefix-materialization logic).

Byte-exactness suites green (both paths), full benchmark unchanged within
noise (NFD Russian 0.33, bert English 0.12, spm Russian 0.35).
…ncode README

The Precompiled wrapper dropped spm_precompiled's inherent `from(&[u8])`
constructor, breaking the Python/Node bindings (and every wheel job) through
the tokenizers crate's re-export — restored as a delegation. The README drift
predates this PR: cargo-readme strips `# `-hidden doctest lines, so README.md
must not carry them.
…C and the feature only feeds the bench comparator (windows clippy runs --all-features)
Drop the ~160 lines of reimplemented legacy normalizers: the baseline is now
the real released crate (tokenizers-release, bench-baseline feature) driven
through its NormalizedString path — what users actually ran. All 99 cells
byte-exact. Highlights vs 0.23.1: bert 41-64 ns/B -> 0.12-6.5 (10-345x),
forms up to 488x, spm 4-60x; worst cell on the board is strip hindi at 3.3x.
run: cargo bench -p tk-encode --bench normalize --features bench-baseline
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants