Lesson 25 / 28
Monitoring, Detection and Logging
Watch for abuse in production without creating a privacy problem.
Signals worth watching
In production, log enough to investigate and detect: requests per user, blocked or flagged inputs, canary-token hits (prompt leakage), policy denials and approval outcomes, tool calls with arguments (redacted), refusals, validation failures, token and cost per user and feature, and latency and error rates. Alert on anomalies: bursts of injection-like inputs, a user repeatedly hitting denials, a sudden jump in cost or tool calls, tools called in unusual sequences, new outbound hosts contacted, or outputs containing secrets or personal data. Balance this with privacy: redact sensitive values, restrict who can read logs, apply retention limits and treat logs as sensitive data. Detection does not stop an attack by itself, so connect it to response: the ability to block a user, disable a tool or feature, and roll back a prompt or model.
Alerts worth having
Thresholds are examples; tune to your traffic.
Signal Example alert rule
canary token appears in an output any hit -> page security, block the response
>= 5 policy denials by one user in 10 min flag the account for review
tool calls per request > 3x normal investigate possible loop or injection
cost per hour > 3x the 7-day baseline alert + automatic spend cap
new outbound hostname from a tool alert; require approval before allowing
secret/PII pattern in model output block, redact, alertAlert on cost spikes
A sudden jump in spend is often the first sign of abuse or a runaway loop.
Quick check: A canary token from the system prompt appears in a reply. What does it indicate?
- The model is faster
- Everything is fine
- The system prompt was leaked and the response should be blocked and reviewed
- The user is a developer
Answer
The system prompt was leaked and the response should be blocked and reviewed — Canaries turn silent leakage into a detectable event.