Public benchmarks are standard test sets used to compare language models: knowledge questions, maths problems, coding tasks, reasoning puzzles and human preference ratings.
What Benchmarks Are Good For
- Quick comparison of general capability across models.
- Tracking progress over time.
- Identifying strengths, such as coding versus multilingual ability.
Why They Can Mislead
- Contamination: benchmark questions may have appeared in training data, inflating scores.
- Saturation: top models may score so highly that differences become meaningless.
- Narrowness: a benchmark measures one kind of task, often in a stylised format.
- Gaming: models can be tuned to benchmarks rather than real use.
- Preference leaderboards reflect what raters like, which may reward style over accuracy.
Read the Fine Print
Check which version of a benchmark was used, the prompting method, the number of attempts allowed, and who ran the test.
Build Your Own Evaluation
The benchmark that matters most is your own task. A small, representative test set of your real inputs, with clear grading criteria, tells you more than any leaderboard about which model to use.
Combine Signals
Use public benchmarks to shortlist candidate models, then choose with your own evaluation, plus cost, latency and data-handling requirements.