Agents raise AI security stakes: a manipulated model that can only chat produces bad text, while a manipulated agent can take harmful actions.
Least Privilege
Give agents only the tools, data and permissions needed for their task. Use scoped, short-lived credentials rather than broad keys.
Isolation
Run agents in sandboxes with restricted file system and network access. Keep secrets outside the sandbox.
Break Dangerous Combinations
The highest risk arises when an agent can access private data, process untrusted content and communicate externally at once. Remove one where possible.
Human Approval
Require confirmation for irreversible, financial or external actions, with clear descriptions of what will happen.
Validate Actions
Check tool arguments against policies in code — not just instructions in the prompt.
Monitor
Log every action; alert on unusual patterns such as bulk data access or unexpected destinations.
Identity
Give each agent its own identity, so its actions are distinguishable from users' and can be audited and revoked.
Test
Red-team agents with injection attempts through every content source they read.