Lesson 1 of 4
Why "it looked good" isn't enough
From vibes to measurable success criteria.
12 min 3-question quiz 3 guides to read next
LLM outputs vary from run to run, work well on typical inputs and fail on unusual ones, and change whenever you edit a prompt or the provider updates a model. Trying a handful of examples by hand can't reveal any of that.
Evaluation starts with success criteria that are:
- Specific: not "good summaries" but "summaries under 100 words that include every figure mentioned in the source and contain no facts that aren't in it".
- Measurable: you can check them automatically, or with a clear rubric a person can apply consistently.
- Tied to the use case: a support bot might be judged on correct answers, correct refusals, tone and whether it escalates when it should.
Split criteria into:
- Must-haves — failures here are serious (leaking personal data, giving a wrong price, missing a safety warning). Aim for zero.
- Quality measures — scores you want to raise over time (helpfulness, concision, style).
Write these down before you tune a single prompt. They decide what "better" means for everything that follows.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Evaluating LLM Outputs
How to measure the quality of language model outputs with test sets, code checks, human review and model-based grading.
2 min read
-
Measuring LLM Quality With Benchmarks
What public benchmarks measure, why leaderboard scores can mislead, and how to complement them with your own tests.
1 min read
-
Defining Good Metrics and KPIs
How to choose metrics that reflect real goals, define them precisely and avoid the traps of vanity metrics and gaming.
2 min read