AI products are often sold with impressive demos. A structured evaluation protects you from expensive surprises.
Accuracy Claims
- Measured on what? Ask how accuracy was measured, on which data and against which baseline.
- Which metric? Accuracy alone can hide poor performance on rare but important cases; ask for precision, recall or error rates by category.
- Across groups? Does performance hold for all the people or situations it will affect?
Test on Your Own Data
Insist on a pilot using a representative sample of your real data, scored against answers agreed by your own experts. Include awkward cases. Vendor examples tell you little about your situation.
Data Handling
- Where is your data processed and stored, and for how long?
- Is it used to train the vendor's models?
- What security certifications and contractual commitments do they offer?
Transparency and Control
Can you see why the system made a decision? Can people override it? What happens when the vendor updates its model — will you be told, and can you re-test?
Total Cost
Look beyond licence fees: integration work, usage-based charges at your real volume, human review, monitoring and exit costs if you switch.
Red Flags
Refusal to allow a pilot on your data; accuracy claims without methodology; "no bias" guarantees; and unclear answers about data use.