BPE vs SentencePiece vs tiktoken: how tokenizers differ

BPE is a tokenization algorithm. SentencePiece and tiktoken are implementations with different training, preprocessing, and vocabulary conventions. SentencePiece supports BPE and unigram models; tiktoken uses byte-level BPE encodings. Their token IDs are not interchangeable.

Engineers estimating context use or processing text must use the tokenizer and prompt template that belong to the exact checkpoint or API model.

Separate algorithm from implementation

Byte Pair Encoding learns repeated merges of adjacent symbols. Character-based BPE begins with characters; byte-level versions begin with bytes. The learned merges and vocabulary determine how a string splits. A familiar word can be one token in one vocabulary and several in another.

NameWhat it specifiesWhat still depends on the model
BPEA merge-based algorithmInitial symbols, preprocessing, merges, vocabulary
SentencePieceA raw-text tokenizer implementation supporting BPE and unigramModel type, normalization, vocabulary, byte fallback
tiktokenAn implementation of byte-level BPE encodingsEncoding, vocabulary, special-token rules

SentencePiece handles raw Unicode text and represents whitespace with its ▁ marker. Byte fallback is optional. tiktoken supplies named encodings and byte-level handling.

Tokenizer families can also change within a model brand. Mistral’s documentation distinguishes earlier SentencePiece tokenizers from Tekken, which uses tiktoken-style BPE. Load the checkpoint’s tokenizer rather than inferring it from the brand.

Measure representative text

For each language and content type you serve, record the token count after applying the actual prompt template. Include role markers, special tokens, tools, and any extra serialized instructions. The visible user text alone may undercount the request.

For comparing tokenizers on a fixed sample, define a unit such as words or characters and divide tokens by that unit count. Keep normalization and the sample fixed. Whitespace-defined “words” are not a comparable linguistic unit for every language, so state the measurement rule.

Vocabulary size alone does not establish token efficiency across different training corpora. A count produced by another model’s tokenizer also does not prove that your request fits the selected model’s context limit.

Engineering Guide: tokenization gives vocabulary examples and explains tokenizer fertility.