Lesson 18 / 25
What to Log and Watch
Capture traces, user feedback and safety events in a way that respects privacy.
Make behaviour visible
For every request record a trace: model and version, prompt template version, retrieved document IDs, tool calls, token counts, latency, cost, the safety decisions made, and the user's feedback (thumbs, edits, escalations). Redact personal data and limit who can read logs. From these you can answer "why did it say that?", find new failure patterns to add to your evaluation set, and notice when quality, safety or spend drifts.
Watch what really happens
Quality, safety and cost drift after launch, so collect signals and set alerts.
One trace record
Log it as a single JSON line so tools can search and chart it. No raw personal data.
{"request_id":"r-5521","model":"model-x","prompt_version":"v14",
"docs":["kb-102","kb-377"],"tools":[],"input_tokens":1480,"output_tokens":212,
"latency_s":1.9,"cost_usd":0.0076,"moderation":"pass","grounded":true,
"feedback":"thumbs_down","pii_masked":2}Version everything
Prompts, retrieval settings and models change behaviour. Putting version labels in every trace lets you link a drop in quality to the exact change that caused it.
Quick check: Why record the prompt version in every trace?
- To link a change in behaviour to the exact version that caused it
- Because logs must be long
- To hide the prompt
- It reduces costs by itself
Answer
To link a change in behaviour to the exact version that caused it — Version labels make regressions traceable to a specific change.