AI costs can grow quickly. Systematic optimisation keeps them sustainable.
Measure First
Track cost by model, feature, team and customer. Find the biggest contributors before optimising.
For LLM APIs
- Use smaller models for simpler tasks, with routing.
- Prompt caching for repeated prefixes.
- Batch APIs for non-urgent work.
- Shorter prompts and output limits.
- Cache responses for repeated questions.
For Self-Hosted Models
- Quantisation and efficient serving frameworks.
- High GPU utilisation through batching.
- Autoscaling to demand.
- Spot capacity for flexible workloads.
For Training
- Start from pretrained models.
- Smaller experiments before large runs.
- Stop unpromising runs early.
Quality Guardrails
Measure quality alongside cost; savings that degrade user experience are false economies.
Review Regularly
Model prices and capabilities change quickly. Re-evaluate options periodically.