Skip to content

LLMOps: Running LLM Applications in Production

The operational practices specific to language model applications: prompt versioning, evaluation, observability, cost and safety.

Editorial team 2 min read

Applications built on large language models need operational practices of their own, often called LLMOps.

What's Different

  • Behaviour depends on prompts, retrieval and tools, not just a trained model.
  • Outputs are open-ended text, harder to evaluate automatically.
  • Models are often third-party services that change over time.
  • Costs are per token and can spike.
  • New risks: prompt injection, hallucination and data leakage.

Core Practices

  • Version prompts and configuration — system prompts, few-shot examples, model names, sampling settings, retrieval parameters and tool definitions — in source control.
  • Evaluation sets run automatically on every change and before upgrading models.
  • Observability: log inputs, outputs, retrieved documents, tool calls, latency, token counts and costs for each request, with privacy safeguards.
  • Guardrails: input and output checks, permission limits for agents.
  • Cost controls: budgets, alerts, rate limits and routing to cheaper models.
  • Human feedback: capture user ratings and corrections, and review samples regularly.

Handling Model Updates

Pin model versions where possible, test new versions against your evaluation set before switching, and keep a fallback.

Incident Readiness

Be able to disable features, switch models or roll back prompts quickly when problems appear.

Measure Business Outcomes

Beyond quality scores, track whether the application actually saves time, resolves issues or improves satisfaction.

More in MLOps & deployment

All MLOps & deployment guides →
MLOps & deployment Guide · 2 min

What Is MLOps?

The practices that take machine learning from notebook to reliable production: versioning, automation, deployment and monitoring.

MLOps & deployment 2 min read 26 Dec 2025

MLOps & deployment Guide · 2 min

Deploying Machine Learning Models

Batch scoring, real-time APIs, streaming and on-device inference: choosing how predictions reach users, and deploying safely.

MLOps & deployment 2 min read 25 Dec 2025

MLOps & deployment Guide · 2 min

Monitoring Machine Learning Models in Production

What to monitor after deployment — data, predictions, outcomes and operations — and how to respond when things change.

MLOps & deployment 2 min read 24 Dec 2025

MLOps & deployment Guide · 2 min

Model Registries and Versioning

Why every production model needs a version, metadata and lineage, and how a model registry manages promotion and rollback.

MLOps & deployment 2 min read 23 Dec 2025