Attacks on AI systems often look like ordinary use. Good monitoring helps spot them.
What to Log
- Prompts and responses, with redaction of sensitive data.
- Retrieved documents and their sources.
- Tool calls, arguments and results.
- User identity, session and rate information.
- Safety classifier results and refusals.
What to Detect
- Known injection and jailbreak patterns.
- Repeated refusals or probing from one user.
- System prompt or secret disclosure in outputs.
- Personal data in outputs.
- Unusual tool use: bulk data access, unexpected destinations.
- Spikes in token use or cost.
- High-volume systematic querying, suggesting extraction.
Alerting and Response
Route alerts to security teams, with enough context to investigate. Define actions: block the user, disable a tool, roll back a change.
Privacy
Logs of AI interactions can be very sensitive. Limit access, set retention periods and tell users what's logged.
Integrate
Feed AI logs into existing security monitoring systems so analysts see AI events alongside other activity.