LLM costs scale with the number of tokens processed. A few design choices can reduce them substantially without hurting quality.
Use the Smallest Model That Works
Test cheaper models on your evaluation set. Classification, extraction and routing often don't need the largest model.
Route Requests
Send simple requests to a small model and hard ones to a larger model. A cheap classifier or rules can decide.
Trim Prompts
- Remove unnecessary instructions and boilerplate.
- Retrieve only the most relevant chunks instead of whole documents.
- Summarise long conversation histories.
Control Output Length
Output tokens usually cost more than input tokens. Ask for concise answers and set maximum output lengths.
Cache
- Prompt caching: many providers discount repeated prompt prefixes such as long system prompts and shared documents; put stable content first.
- Response caching: store answers to identical or near-identical requests.
Batch
For work that isn't time-sensitive, batch APIs often offer lower prices in exchange for slower turnaround.
Measure
Track tokens and cost per request, per feature and per user. Set budgets and alerts. Watch for runaway agent loops.
Balance Against Quality
Re-run your evaluation after every cost optimisation. A cheaper system that fails users costs more in the end.