Skip to content

Error Handling and Recovery in Agents

How agent harnesses cope when tools fail, models misbehave or tasks go wrong — retries, feedback and graceful stopping.

Editorial team 1 min read

Agents run into errors constantly: commands fail, APIs time out, files are missing, and models occasionally produce malformed tool calls. Robust harnesses expect this.

Return Errors to the Model

Tool errors should come back as informative observations. A clear message lets the model adjust its approach.

Retry Transient Failures

Network timeouts and rate limits can be retried automatically with backoff, without involving the model.

Validate Tool Calls

Check arguments against schemas before execution. Return validation errors so the model can correct them.

Detect Loops

Agents sometimes repeat the same failing action. Detect repetition and intervene: add a hint, change strategy, or stop.

Checkpoints

Save state at key points so work can resume after crashes, and changes can be rolled back if the agent makes a mess — version control is ideal for code.

Know When to Stop

Set limits on steps, time and cost. When the agent is stuck, it should report what it tried and what blocked it rather than continuing indefinitely or claiming success.

Honest Reporting

Instruct agents to report failures plainly. Evaluate for false claims of success — it's a common and costly failure mode.

More in Agent harnesses

All Agent harnesses guides →
Agent harnesses Guide · 2 min

What Is an Agent Harness?

The software around a language model that turns it into an agent: the loop, tools, context, permissions and memory.

Agent harnesses 2 min read 27 Sep 2025

Agent harnesses Guide · 1 min

The Agent Loop Explained

The core cycle every agent runs: think, call a tool, observe the result, repeat — and how the loop knows when to stop.

Agent harnesses 1 min read 26 Sep 2025

Agent harnesses Guide · 1 min

Designing Tools for AI Agents

How to write tools agents use well: clear names, precise descriptions, sensible inputs and informative outputs.

Agent harnesses 1 min read 25 Sep 2025

Agent harnesses Guide · 1 min

Context Management in Agent Harnesses

How agents stay effective over long tasks: what to keep in context, what to summarise, and what to store outside.

Agent harnesses 1 min read 24 Sep 2025