Language models don't read characters or words directly. They read tokens — chunks of text from a fixed vocabulary.
How It Works
Tokenisers such as byte-pair encoding build a vocabulary of common character sequences. Frequent words become single tokens; rare words split into pieces. "Unbelievable" might become "un", "believ", "able".
Why It Matters
- Cost: APIs charge per token.
- Limits: context windows and output limits are measured in tokens.
- Languages: many non-English languages need more tokens per word, raising cost.
- Quirks: models can struggle with tasks that depend on characters within tokens — counting letters, reversing words, some arithmetic.
Rough Rules
For English, a token averages roughly three-quarters of a word, but this varies with the tokeniser and content. Code, numbers and unusual text use more tokens.
Counting Tokens
Providers offer tokeniser libraries or token-counting endpoints. Use them for accurate cost and limit calculations.
Different Models, Different Tokenisers
Token counts for the same text differ between model families, so compare costs using each model's own tokeniser.