Lesson 25 / 29

Safety Essentials Every LLM Feature Needs

Apply the core controls; go deeper in the security and AI-safety courses.

A short list that prevents most trouble

Every LLM feature, however small, should have these controls. Treat model output as untrusted (escape, parameterise, validate). Treat retrieved and user content as untrusted instructions: fence and label it, never let it grant permissions. Least privilege for tools and data, with per-user authorisation and human approval for risky actions. No secrets in prompts or logs; redact personal data. Rate limits and spend caps. Content safety appropriate to your audience (moderation, refusal behaviour, escalation paths for sensitive topics such as health or self-harm). Honest UX: tell users they are talking to an AI, show sources, express uncertainty, and offer a human. Red-team tests in CI. The dedicated courses on LLM application security and AI safety cover each of these in depth; this checklist is the minimum bar to clear before launch.

Responsible by design, documented, shared

Security basics, privacy, documentation and clear roles keep LLM features trustworthy as they and the team grow.

Three needs: safe, accountable, documented.
Figure 7.1 — Safe, accountable and documented.

A pre-launch safety checklist

Tick every box or write down why not.

[ ] model output escaped / parameterised / schema-validated before any use
[ ] retrieved + user content fenced, labelled as data; cannot grant permissions
[ ] tools least-privilege; per-user authorisation; approvals for risky actions; default-deny policy
[ ] no secrets in prompts/logs; PII minimised and redacted
[ ] rate limits, token/step caps, spend cap + alert
[ ] content-safety behaviour tested (refusals, sensitive topics, escalation to a human)
[ ] UX: AI disclosure, sources shown, uncertainty expressed, human handoff available
[ ] red-team suite in CI; incident runbook + kill switches rehearsed

Rehearse the kill switch

Disable a tool or feature in staging to confirm the app degrades gracefully.

Quick check: Which is part of an honest UX for an LLM feature?

  • Hiding all uncertainty
  • Pretending to be a human
  • Telling users it is an AI, showing sources and offering a human
  • Removing citations
Answer

Telling users it is an AI, showing sources and offering a human — Transparency builds appropriate trust and reduces harm from wrong answers.