Agents combine a model that can be manipulated with tools that can act. That makes security central to harness design.
Main Threats
- Prompt injection: instructions hidden in web pages, emails, documents, code comments or tool results that hijack the agent.
- Data exfiltration: tricking the agent into sending private data to an attacker, for example via a URL or message.
- Excessive permissions: agents with broad access that turn small mistakes into large incidents.
- Malicious tools and integrations: untrusted plugins or servers that misbehave or return manipulative content.
- Secrets exposure: credentials in files, environment variables or logs that the agent reads and leaks.
The Dangerous Combination
Risk is highest when an agent has access to private data, reads untrusted content, and can communicate externally. Remove at least one of these where possible.
Defences
- Least privilege and scoped, short-lived credentials.
- Sandboxing with restricted network access.
- Approval for consequential actions.
- Treat all tool output and retrieved content as untrusted data.
- Vet tools and integrations before installing.
- Monitor and log actions.
Test Adversarially
Red-team your agents with injection attempts through every input channel before deployment.