BPE vs SentencePiece vs tiktoken: how tokenizers differ
BPE is a tokenization algorithm. SentencePiece and tiktoken are implementations with different training, preprocessing, and vocabulary conventions. SentencePiece supports BPE and unigram models; tiktoken uses byte-level BPE encodings. Their token IDs are not interchangeable.
Engineers estimating context use or processing text must use the tokenizer and prompt template that belong to the exact checkpoint or API model.
Separate algorithm from implementation
Byte Pair Encoding learns repeated merges of adjacent symbols. Character-based BPE begins with characters; byte-level versions begin with bytes. The learned merges and vocabulary determine how a string splits. A familiar word can be one token in one vocabulary and several in another.
| Name | What it specifies | What still depends on the model |
|---|---|---|
| BPE | A merge-based algorithm | Initial symbols, preprocessing, merges, vocabulary |
| SentencePiece | A raw-text tokenizer implementation supporting BPE and unigram | Model type, normalization, vocabulary, byte fallback |
| tiktoken | An implementation of byte-level BPE encodings | Encoding, vocabulary, special-token rules |
SentencePiece handles raw Unicode text and represents whitespace with its ▁ marker. Byte fallback is optional. tiktoken supplies named encodings and byte-level handling.
Tokenizer families can also change within a model brand. Mistral’s documentation distinguishes earlier SentencePiece tokenizers from Tekken, which uses tiktoken-style BPE. Load the checkpoint’s tokenizer rather than inferring it from the brand.
Measure representative text
For each language and content type you serve, record the token count after applying the actual prompt template. Include role markers, special tokens, tools, and any extra serialized instructions. The visible user text alone may undercount the request.
For comparing tokenizers on a fixed sample, define a unit such as words or characters and divide tokens by that unit count. Keep normalization and the sample fixed. Whitespace-defined “words” are not a comparable linguistic unit for every language, so state the measurement rule.
Vocabulary size alone does not establish token efficiency across different training corpora. A count produced by another model’s tokenizer also does not prove that your request fits the selected model’s context limit.
Engineering Guide: tokenization gives vocabulary examples and explains tokenizer fertility.