When a user reports a wrong answer, you need to see exactly what happened. Good logging makes RAG debuggable.
What to Log per Request
- The original question and any rewritten queries.
- Retrieved chunk IDs, scores and sources — before and after re-ranking.
- The final prompt (or a reference to its template version).
- The model, settings and response.
- Citations produced.
- Latency per stage and token counts.
- User feedback, if given.
Dashboards
Track answer volume, feedback ratings, "couldn't find" rates, latency percentiles, cost per request and the most common questions.
Finding Problems
- Questions with negative feedback.
- Answers with no citations.
- Retrieval with low top scores (likely content gaps).
- Spikes in latency or cost.
Privacy
Logs may contain personal and confidential information. Restrict access, redact where possible and set retention limits.
Sampling for Review
Regularly review a random sample of interactions, not just flagged ones, to catch silent failures.
Close the Loop
Turn discovered failures into test cases, and track whether fixes improve metrics over time.