An evaluation harness is software that runs a model against a set of test cases and scores the results consistently. (It's related to, but distinct from, an agent harness.)
Components
- Datasets: test cases with inputs and expected outputs or grading criteria.
- Prompt formatting: turning each case into the exact prompt, including any few-shot examples.
- Model interface: calling local or hosted models with fixed settings.
- Scoring: exact match, multiple-choice accuracy, unit tests for code, similarity metrics, or model-based grading.
- Reporting: aggregate scores, breakdowns by category, and individual failures.
Why Consistency Matters
Small changes in prompt format, number of examples or answer extraction can change scores noticeably. A shared harness makes comparisons between models fair.
Open-Source Harnesses
Several open-source harnesses run standard academic benchmarks. They're useful for comparing models on common tasks.
Custom Evaluation
Standard benchmarks rarely reflect your application. Build your own test sets from real tasks and run them through the same kind of pipeline.
Pitfalls
- Benchmark contamination: test data seen in training.
- Overfitting to the benchmark rather than the real task.
- Ignoring variance between runs.