# What to Log and Watch — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/mon-observability

> Capture traces, user feedback and safety events in a way that respects privacy.

## Make behaviour visible

For every request record a **trace**: model and version, prompt template version, retrieved document IDs, tool calls, token counts, latency, cost, the safety decisions made, and the user's **feedback** (thumbs, edits, escalations). Redact personal data and limit who can read logs. From these you can answer "why did it say that?", find new failure patterns to add to your evaluation set, and notice when quality, safety or spend drifts.

## Watch what really happens

Quality, safety and cost drift after launch, so collect signals and set alerts.

![Three signals: quality, safety, spend.](assets/figures/ai-safety/section-6-map.svg) — Figure 6.1 — Quality, safety and spend signals.

## One trace record

Log it as a single JSON line so tools can search and chart it. No raw personal data.

```json
{"request_id":"r-5521","model":"model-x","prompt_version":"v14",
 "docs":["kb-102","kb-377"],"tools":[],"input_tokens":1480,"output_tokens":212,
 "latency_s":1.9,"cost_usd":0.0076,"moderation":"pass","grounded":true,
 "feedback":"thumbs_down","pii_masked":2}
```

## Version everything

Prompts, retrieval settings and models change behaviour. Putting version labels in every trace lets you link a drop in quality to the exact change that caused it.

**Quiz:** Why record the prompt version in every trace?

- [x] To link a change in behaviour to the exact version that caused it
- [ ] Because logs must be long
- [ ] To hide the prompt
- [ ] It reduces costs by itself

*Answer:* To link a change in behaviour to the exact version that caused it. Version labels make regressions traceable to a specific change.
