Skip to content

Evaluating LLM Outputs

How to measure the quality of language model outputs with test sets, code checks, human review and model-based grading.

Editorial team 2 min read

"It looked good in the demo" isn't evaluation. LLM outputs vary, fail on unusual inputs and change when prompts or models change.

Define Success First

Write specific, measurable criteria: summaries under 100 words that include every figure in the source and add no facts not in it. Separate must-haves (safety, correctness) from quality measures (helpfulness, tone).

Build a Test Set

Collect realistic inputs — real user questions, historical documents — plus edge cases, out-of-scope requests and past failures. Start with 50–100 well-chosen examples and grow it.

Three Ways to Grade

  • Code-based checks: valid JSON, required fields, length limits, exact matches, numbers within tolerance. Fast and consistent — use wherever possible.
  • Human review: best for nuanced quality, with a written rubric and more than one reviewer on a sample.
  • Model-based grading: an LLM scores outputs against a rubric. Scales well; validate it against human judgements, grade one criterion at a time, and ask for reasoning before the verdict.

Run It Continuously

Re-run the evaluation whenever the prompt, model, retrieval setup or tools change, and compare with the previous run. Treat regressions on must-haves as release blockers.

Look at Failures

Averages hide important errors. Read individual failures, and add new ones to the test set.

Consider Cost and Speed

A slightly better score at four times the cost or latency may not be the right trade.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026