Subword-tokenization pre-filter for byte-level compressors, and the harness that tested it. +7-9.6% over raw bytes on LZMA2; the gap widens with corpus size then plateaus at the dictionary window. 452 matrix cells on PG-19 (64MB-4GB), all sha256 round-trip verified.
benchmark compression reproducible-research data-compression zstd lzma xz tokenization bpe byte-pair-encoding tiktoken pg19
-
Updated
Aug 21, 2026 - Python