Add a ParityBpeTrainer example and list it in the trainers API docs - #2217
Open
cimeister wants to merge 1 commit into
Open
Add a ParityBpeTrainer example and list it in the trainers API docs#2217cimeister wants to merge 1 commit into
cimeister wants to merge 1 commit into
Conversation
ParityBpeTrainer shipped without a usage example or an API docs entry. Its docstring assumed local per-language files that users are unlikely to have, and did not document the `ratio` argument. Add examples/train_parity_bpe.py, which pulls per-language training text from wikimedia/wikipedia and a parallel dev set from openlanguagedata/flores_plus, then demonstrates both balancing modes. List ParityBpeTrainer in docs/source-doc-builder/api/trainers.mdx. Document `ratio` in more detail: it is a relative target, the trainer selects the language with the lowest compression_rate / ratio, and compression is counted in the units the pre-tokenizer emits, which for ByteLevel means bytes. Equal ratios therefore do not give equal tokenization across scripts, since Devanagari takes about 2.5x the bytes of Latin script for the same content. The example derives its ratios from mean FLORES+ bytes per sentence for that reason. Also note that balancing needs either a parallel dev set or ratios, and that the dev set takes precedence when both are passed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ParityBpeTrainer was added without a usage example or an API docs entry. Its docstring assumed local per-language files that users are unlikely to have, and did not document the
ratioargument.This PR:
ratioparameter in more detail: it is a relative target, the trainer selects the language with the lowest compression_rate / ratio, and compression is counted in the units the pre-tokenizer emits, which for ByteLevel means bytes. Equal ratios therefore do not give equal tokenization across scripts, since Devanagari takes about 2.5x the bytes of Latin script for the same content. The example computes its ratios from mean FLORES+ bytes per sentence for that reason.