Model output often flows into other systems: web pages, databases, shells, emails. Treating it as trusted is a common and serious mistake.
Risks
- Cross-site scripting: model output rendered as HTML can include malicious scripts.
- SQL injection: generated queries executed without safeguards.
- Command injection: output passed to shells.
- Server-side request forgery: model-chosen URLs fetched by servers.
- Data exfiltration: output containing links or images that send data to attackers when rendered.
Defences
- Encode output appropriately for its context — HTML-escape text displayed in web pages.
- Sanitise Markdown and HTML before rendering; restrict links and images to trusted domains.
- Parameterised queries and read-only database accounts for generated SQL.
- Allow-lists for commands, URLs and file paths.
- Schema validation for structured output.
- Sandbox execution of generated code.
Principle
Apply the same controls you would to user input. The model may have been influenced by an attacker, directly or through content it read.