Agents may call MCP tools many times per task. Slow or flaky servers make agents slow and flaky.
Latency
- Cache results that don't change often.
- Keep connections to back-end systems warm.
- Return concise results; large payloads take time to transfer and process.
Long Operations
For slow tasks, send progress notifications where supported, or start the work and return a handle the model can check later.
Timeouts
Set timeouts on back-end calls and return clear errors when they're exceeded, rather than hanging.
Cancellation
The protocol supports cancelling in-progress requests. Honour cancellations to avoid wasted work.
Graceful Failure
When a back end is down, return an informative error so the model can tell the user or try another approach.
Rate Limits
Protect back ends with limits, and return messages explaining when to retry.
Monitoring
Track latency percentiles, error rates and call volumes per tool. Unusual spikes can signal an agent stuck in a loop.
Load Testing
Simulate agent traffic patterns — bursts of rapid calls — before launch.