Skip to content

Reducing the Cost of LLM Applications

Practical ways to cut token costs: model choice, prompt size, caching, batching and routing.

Editorial team 2 min read

LLM costs scale with the number of tokens processed. A few design choices can reduce them substantially without hurting quality.

Use the Smallest Model That Works

Test cheaper models on your evaluation set. Classification, extraction and routing often don't need the largest model.

Route Requests

Send simple requests to a small model and hard ones to a larger model. A cheap classifier or rules can decide.

Trim Prompts

  • Remove unnecessary instructions and boilerplate.
  • Retrieve only the most relevant chunks instead of whole documents.
  • Summarise long conversation histories.

Control Output Length

Output tokens usually cost more than input tokens. Ask for concise answers and set maximum output lengths.

Cache

  • Prompt caching: many providers discount repeated prompt prefixes such as long system prompts and shared documents; put stable content first.
  • Response caching: store answers to identical or near-identical requests.

Batch

For work that isn't time-sensitive, batch APIs often offer lower prices in exchange for slower turnaround.

Measure

Track tokens and cost per request, per feature and per user. Set budgets and alerts. Watch for runaway agent loops.

Balance Against Quality

Re-run your evaluation after every cost optimisation. A cheaper system that fails users costs more in the end.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026