Skip to content

Security Threats to Agent Harnesses

The main ways agents can be attacked or go wrong — prompt injection, data exfiltration, over-broad permissions — and defences.

Editorial team 1 min read

Agents combine a model that can be manipulated with tools that can act. That makes security central to harness design.

Main Threats

  • Prompt injection: instructions hidden in web pages, emails, documents, code comments or tool results that hijack the agent.
  • Data exfiltration: tricking the agent into sending private data to an attacker, for example via a URL or message.
  • Excessive permissions: agents with broad access that turn small mistakes into large incidents.
  • Malicious tools and integrations: untrusted plugins or servers that misbehave or return manipulative content.
  • Secrets exposure: credentials in files, environment variables or logs that the agent reads and leaks.

The Dangerous Combination

Risk is highest when an agent has access to private data, reads untrusted content, and can communicate externally. Remove at least one of these where possible.

Defences

  • Least privilege and scoped, short-lived credentials.
  • Sandboxing with restricted network access.
  • Approval for consequential actions.
  • Treat all tool output and retrieved content as untrusted data.
  • Vet tools and integrations before installing.
  • Monitor and log actions.

Test Adversarially

Red-team your agents with injection attempts through every input channel before deployment.

More in Agent harnesses

All Agent harnesses guides →
Agent harnesses Guide · 2 min

What Is an Agent Harness?

The software around a language model that turns it into an agent: the loop, tools, context, permissions and memory.

Agent harnesses 2 min read 27 Sep 2025

Agent harnesses Guide · 1 min

The Agent Loop Explained

The core cycle every agent runs: think, call a tool, observe the result, repeat — and how the loop knows when to stop.

Agent harnesses 1 min read 26 Sep 2025

Agent harnesses Guide · 1 min

Designing Tools for AI Agents

How to write tools agents use well: clear names, precise descriptions, sensible inputs and informative outputs.

Agent harnesses 1 min read 25 Sep 2025

Agent harnesses Guide · 1 min

Context Management in Agent Harnesses

How agents stay effective over long tasks: what to keep in context, what to summarise, and what to store outside.

Agent harnesses 1 min read 24 Sep 2025