"It looked good in the demo" isn't evaluation. LLM outputs vary, fail on unusual inputs and change when prompts or models change.
Define Success First
Write specific, measurable criteria: summaries under 100 words that include every figure in the source and add no facts not in it. Separate must-haves (safety, correctness) from quality measures (helpfulness, tone).
Build a Test Set
Collect realistic inputs — real user questions, historical documents — plus edge cases, out-of-scope requests and past failures. Start with 50–100 well-chosen examples and grow it.
Three Ways to Grade
- Code-based checks: valid JSON, required fields, length limits, exact matches, numbers within tolerance. Fast and consistent — use wherever possible.
- Human review: best for nuanced quality, with a written rubric and more than one reviewer on a sample.
- Model-based grading: an LLM scores outputs against a rubric. Scales well; validate it against human judgements, grade one criterion at a time, and ask for reasoning before the verdict.
Run It Continuously
Re-run the evaluation whenever the prompt, model, retrieval setup or tools change, and compare with the previous run. Treat regressions on must-haves as release blockers.
Look at Failures
Averages hide important errors. Read individual failures, and add new ones to the test set.
Consider Cost and Speed
A slightly better score at four times the cost or latency may not be the right trade.