Agents make many model calls with growing context, so costs and time add up quickly.
Where Cost Comes From
- Repeatedly sending long context — instructions, tools, history — on every step.
- Large tool outputs.
- Many steps on complex tasks.
- Sub-agents multiplying token usage.
Reducing Cost
- Prompt caching: reuse the stable prefix of each request (instructions, tool definitions) where providers support it.
- Trim tool output: return only what's needed.
- Right-size models: use smaller, faster models for simple sub-tasks.
- Compaction: summarise history rather than resending everything.
- Budgets: cap steps, tokens and spend per task.
Reducing Latency
- Run independent tool calls in parallel.
- Stream progress to users.
- Use faster models where quality allows.
- Avoid unnecessary steps with better tools and instructions.
Measure Per Task
Track cost and time per completed task, not per call. A more capable model that finishes in fewer steps can be cheaper overall.