Guardrails are the checks around a language model that keep an application safe, accurate and within its intended scope.
Input Guardrails
- Scope checks: detect requests outside what the application should handle.
- Sensitive data detection: spot personal or confidential data before it reaches the model.
- Abuse and injection detection: flag attempts to manipulate the model.
- Rate limits and size limits.
Output Guardrails
- Format validation: ensure structured output parses and matches a schema.
- Grounding checks: verify claims are supported by retrieved sources.
- Content moderation: screen for harmful or inappropriate content.
- Data leakage checks: prevent secrets or other users' data appearing in responses.
- Business rules: prices, discounts or commitments the assistant may not make.
Action Guardrails for Agents
Restrict which tools can be called, validate arguments, require confirmation for irreversible actions, and cap steps and spending.
Implementing Them
Guardrails can be simple code rules, classifiers, a second model acting as a checker, or dedicated guardrail frameworks. Deterministic checks in code are the most reliable; use model-based checks for nuanced judgements.
Balance
Overly strict guardrails frustrate users and block legitimate requests. Evaluate them like any other component: measure both what they catch and what they wrongly block.
Log and Review
Record guardrail triggers, review samples regularly and refine the rules as real usage reveals gaps.