Trained models embody expensive data, compute and expertise. Attackers may try to steal them.
Direct Theft
Stealing model weights through compromised servers, insider access, leaked credentials or insecure storage. For frontier models, weight security is a major concern.
Extraction Through Queries
Querying a model many times and training a copy on the responses — sometimes called distillation or model stealing. The copy may approximate the original's behaviour for a fraction of the cost.
Risks
- Loss of competitive advantage.
- Stolen models used without safety safeguards.
- Copies used to develop attacks against the original.
Protecting Weights
- Strict access control and multi-person approval.
- Encryption at rest and in transit.
- Isolated, monitored infrastructure.
- Insider-risk programmes.
Limiting Extraction
- Rate limits and quotas per user.
- Monitoring for high-volume, systematic querying.
- Terms of service prohibiting training on outputs.
- Returning only necessary outputs, not full probability distributions.
Balance
Protections must be weighed against legitimate use; excessive restrictions frustrate real customers.