Language models don't read words the way we do. They read tokens, and they can only consider so many at a time.
What Is a Token?
A tokenizer splits text into pieces from a fixed vocabulary. Common words are often a single token; rarer words are split into several pieces; punctuation and spaces count too. As a rough rule for English, a token is about three-quarters of a word, so 1,000 tokens is roughly 750 words. Other languages and code can use more tokens per word.
Why Tokens Matter
- Cost: API pricing for language models is usually per million input and output tokens.
- Speed: generating output takes time per token.
- Limits: every model has a maximum number of tokens it can handle.
The Context Window
The context window is the maximum number of tokens a model can process in one request, including your instructions, any documents, the conversation so far and its own reply. Anything outside the window simply doesn't exist for the model.
Working Within the Limit
- Put only relevant material in the prompt; long, noisy context can lower quality.
- Summarise earlier parts of long conversations.
- For large document collections, use retrieval: search for the relevant passages and include just those.
- Place key instructions clearly, and consider repeating critical constraints near the end of long prompts.
Counting Tokens
Model providers offer tokenizer tools or APIs to count tokens. Count real samples of your data rather than guessing, especially for non-English text.