AI inference is expensive. Attackers — or buggy clients — can exploit that.
Attack Types
- Volume attacks: flooding an AI endpoint with requests.
- Expensive inputs: very long prompts or files that consume many tokens.
- Output amplification: prompts designed to produce extremely long responses.
- Agent loops: tricking agents into repeating actions indefinitely.
- Resource-heavy tools: triggering costly searches or computations.
- Account abuse: free tiers or stolen API keys used to consume resources.
Defences
- Authentication and per-user rate limits.
- Limits on input size, output length and conversation length.
- Step, time and cost budgets for agents.
- Spending caps and alerts with model providers.
- Caching for repeated requests.
- Anomaly detection on usage patterns.
- Protection of API keys, which are valuable to attackers.
Business Design
Pricing and free tiers should account for abuse. Require verification for higher usage levels.
Monitor
Track cost per user and per feature, and investigate sudden increases quickly.