Lesson 18 / 28

Sensitive Information Disclosure and Redaction

Minimise and redact personal data before it reaches the model or the logs.

Data minimisation is the strongest control

A model can reveal sensitive information it was given: personal data in a prompt can appear in a reply to someone else (via shared context, caching mistakes or retrieval), end up in logs and analytics, or be sent to a third-party provider. Principles: send the minimum (do the model need the customer's full address to summarise a complaint?); redact or pseudonymise direct identifiers (names, phone numbers, emails, ID and card numbers) before the call and restore them afterwards if needed; keep secrets and credentials out of prompts entirely; separate per-user context so one user's data cannot be mixed into another's conversation; apply retention limits to prompts, outputs and logs; and check provider data-use terms. Pattern-based redaction is a safety net with gaps (unusual formats, spelled-out digits, free-text identifiers, other scripts), so combine it with minimisation and access control.

What goes in can come out

Keep sensitive data out of prompts and logs, enforce access in retrieval, and guard the knowledge the model learns from.

Three fronts: disclosure, retrieval, poisoning.
Figure 5.1 — Disclosure, retrieval and poisoning.

Pattern-based redaction and its gaps, run

I ran this with plain Python 3 (standard library only). All attacks here are harmless demonstrations on local data, using no real systems. The email, an Indian-style mobile number (with spaces) and a card number are replaced with labels, but the Hindi name is untouched and a phone number written out in words slips through. Redaction patterns cover structured identifiers only.

import re

PATTERNS = [
    ("EMAIL", re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")),
    ("PHONE", re.compile(r"(?<!\d)(?:\+91[- ]?)?[6-9]\d{4}[ -]?\d{5}(?!\d)")),
    ("CARD",  re.compile(r"(?<!\d)(?:\d[ -]?){13,16}(?!\d)")),
]
def redact(text):
    for label, pat in PATTERNS: text = pat.sub(f"[{label}]", text)
    return text

msg = "Reach me at asha.k@example.com or +91 98765 43210; card 4111 1111 1111 1111. Hindi name: आशा"
print(redact(msg))
print(redact("call nine eight seven six five four three two one zero"))     # spelled-out digits slip through

Output:

Reach me at [EMAIL] or [PHONE]; card [CARD]. Hindi name: आशा
call nine eight seven six five four three two one zero

Do not log full prompts by default

Log metadata and redacted samples; full prompts often contain personal data.

Quick check: What is the strongest control against sensitive data leaking through an LLM app?

  • Not sending sensitive data to the model or logs in the first place
  • Asking the model not to repeat it
  • A longer system prompt
  • Using more tokens
Answer

Not sending sensitive data to the model or logs in the first place — Data that never enters the prompt cannot be repeated.