Lesson 1 / 25

What AI Safety Means for Applications

Define application-level safety as preventing harm to users, the business and third parties.

Safety is a product property

When you build with a language model you are responsible for how the whole application behaves, not only the model. AI safety here means reducing the chance and impact of harm: wrong or invented answers, harmful content, leaks of private data, unfair treatment, misuse by bad actors, and over-trust by users. Model providers add their own safeguards, but they cannot know your users, data or context, so application-level controls are still your job.

Layers of protection

Safe LLM applications combine good prompts, filters, grounded data, limited tools and human oversight.

Four layers: input, model, output, oversight.
Figure 1.1 — Input, model, output and oversight.

A restaurant kitchen

A supplier delivers safe ingredients, yet the restaurant still needs hygiene rules, allergy labels and temperature checks. The supplier's quality does not replace the kitchen's responsibility.

Start with who could be harmed

For each feature, list who could be hurt and how (a student given wrong medical advice, a customer whose data is exposed, a person described unfairly). The list tells you where safeguards matter most.

Quick check: Who is responsible for how an LLM application behaves?

  • Only the model provider
  • Nobody
  • The team that builds and deploys the application
  • Only the end user
Answer

The team that builds and deploys the application — Providers add safeguards, but the application owner controls context, data, tools and deployment.