Chatbots and agents both use language models, but they fail differently and need different evaluation.
Chatbots
Evaluation focuses on individual responses:
- Accuracy and groundedness.
- Helpfulness and tone.
- Safety and refusals.
- Format and length.
Test sets of questions with rubrics or reference answers work well.
Agents
Agents take many steps and change things in the world. Evaluate:
- Task success: did the end state match the goal?
- Efficiency: steps, time and cost.
- Safety: did it avoid forbidden actions and ask for approval when needed?
- Honesty: did it report what it actually did?
- Robustness: does it recover from errors?
Methods for Agents
- Sandboxed environments with automated end-state checks.
- Repeated runs to measure success rates.
- Trajectory review for insight into failures.
Shared Principles
- Use realistic tasks from your domain.
- Combine automated scoring with human review.
- Rerun evaluations whenever models, prompts or tools change.