Applications built on large language models need operational practices of their own, often called LLMOps.
What's Different
- Behaviour depends on prompts, retrieval and tools, not just a trained model.
- Outputs are open-ended text, harder to evaluate automatically.
- Models are often third-party services that change over time.
- Costs are per token and can spike.
- New risks: prompt injection, hallucination and data leakage.
Core Practices
- Version prompts and configuration — system prompts, few-shot examples, model names, sampling settings, retrieval parameters and tool definitions — in source control.
- Evaluation sets run automatically on every change and before upgrading models.
- Observability: log inputs, outputs, retrieved documents, tool calls, latency, token counts and costs for each request, with privacy safeguards.
- Guardrails: input and output checks, permission limits for agents.
- Cost controls: budgets, alerts, rate limits and routing to cheaper models.
- Human feedback: capture user ratings and corrections, and review samples regularly.
Handling Model Updates
Pin model versions where possible, test new versions against your evaluation set before switching, and keep a fallback.
Incident Readiness
Be able to disable features, switch models or roll back prompts quickly when problems appear.
Measure Business Outcomes
Beyond quality scores, track whether the application actually saves time, resolves issues or improves satisfaction.