Lesson 3 / 25

Safety by Design: Layers

Combine input checks, careful prompts, grounded data, limited tools, output checks and human oversight.

No single filter is enough

Each safeguard has gaps, so stack them. Before the model: validate input, detect abuse, strip or mask personal data. In the model call: clear instructions, retrieved evidence, and only the tools the task needs. After the model: check the output (moderation, format, grounding) before showing it. Around it: rate limits, logging, human review for high-stakes actions, and a way for users to report problems.

A guarded request pipeline

Each step can reject or reshape the request. Function names are placeholders for your own code.

def handle(user_msg, user):
    if not rate_limiter.allow(user):          return refuse("Too many requests")
    msg = scrub_pii(user_msg)                 # mask personal data
    if moderation(msg).flagged:               return refuse("Cannot help with that")
    context = retrieve_documents(msg)         # ground the answer
    draft = call_model(msg, context, tools=SAFE_TOOLS)
    if not grounded(draft, context):          return fallback("I could not find that in our documents.")
    if moderation(draft).flagged:             return refuse("Cannot share that")
    log(user, msg, draft)
    return draft

Fail safe, not silent

When a check fails, return a helpful safe response ("I could not verify that") rather than the unchecked draft or a blank error. Users trust systems that admit limits.

Quick check: Why stack several safeguards instead of relying on one?

  • It removes the need for testing
  • One safeguard is illegal
  • Layers make the model faster
  • Each has gaps, so layers catch what others miss
Answer

Each has gaps, so layers catch what others miss — Independent layers reduce the chance that a single miss becomes harm.