Lesson 3 / 25
Safety by Design: Layers
Combine input checks, careful prompts, grounded data, limited tools, output checks and human oversight.
No single filter is enough
Each safeguard has gaps, so stack them. Before the model: validate input, detect abuse, strip or mask personal data. In the model call: clear instructions, retrieved evidence, and only the tools the task needs. After the model: check the output (moderation, format, grounding) before showing it. Around it: rate limits, logging, human review for high-stakes actions, and a way for users to report problems.
A guarded request pipeline
Each step can reject or reshape the request. Function names are placeholders for your own code.
def handle(user_msg, user):
if not rate_limiter.allow(user): return refuse("Too many requests")
msg = scrub_pii(user_msg) # mask personal data
if moderation(msg).flagged: return refuse("Cannot help with that")
context = retrieve_documents(msg) # ground the answer
draft = call_model(msg, context, tools=SAFE_TOOLS)
if not grounded(draft, context): return fallback("I could not find that in our documents.")
if moderation(draft).flagged: return refuse("Cannot share that")
log(user, msg, draft)
return draftFail safe, not silent
When a check fails, return a helpful safe response ("I could not verify that") rather than the unchecked draft or a blank error. Users trust systems that admit limits.
Quick check: Why stack several safeguards instead of relying on one?
- It removes the need for testing
- One safeguard is illegal
- Layers make the model faster
- Each has gaps, so layers catch what others miss
Answer
Each has gaps, so layers catch what others miss — Independent layers reduce the chance that a single miss becomes harm.