Threat modelling asks: what are we building, what can go wrong, what will we do about it, and did we do a good job? It adapts well to AI systems.
Map the System
Draw the data flows: users, prompts, models, retrieval sources, tools, outputs, logs and third parties. Mark trust boundaries — where untrusted content enters.
Identify Threats
For each component and flow, consider:
- Prompt injection, directly and via retrieved content.
- Data leakage in outputs, logs or to providers.
- Excessive permissions for tools and agents.
- Poisoned data or models.
- Unsafe output handling.
- Abuse: spam, fraud, cost exhaustion.
- Traditional threats: authentication, access control, dependencies.
Frameworks like STRIDE, the OWASP LLM Top 10 and MITRE ATLAS help prompt thinking.
Assess and Prioritise
Estimate likelihood and impact, focusing on realistic attackers and your most sensitive assets.
Define Controls
Choose mitigations for high-priority threats, and record accepted risks.
Revisit
Update the model when you add tools, data sources or capabilities — each addition changes the threat landscape.