Lesson 23 / 27

Prompt Injection, Privacy and Bias

Recognise the main security and ethics risks of LLM applications.

Text can carry commands

Prompt injection: instructions hidden in user input or in retrieved content (a web page, an email, a PDF) can override your rules, because the model cannot reliably tell data from instructions. Defences: keep secrets out of prompts, give tools the least privilege, validate and sanitise model output before it reaches shells, SQL or browsers, require human approval for side-effecting actions, and separate trusted instructions from untrusted text. Privacy: do not send personal data to providers without a lawful basis and a data-handling agreement; log carefully and redact. Bias and fairness: models absorb stereotypes from data, so test across groups and languages and monitor outcomes. Follow the laws and policies that apply to you.

Assume the model can be tricked

Design so that even a fully hijacked model cannot do serious harm: narrow tools, read-only access where possible, approvals for anything irreversible.

Quick check: Which design best limits damage from prompt injection?

  • Least-privilege tools plus human approval for risky actions
  • Giving the model admin access
  • Hiding the system prompt only
  • Using a higher temperature
Answer

Least-privilege tools plus human approval for risky actions — Limiting what a compromised model can do bounds the harm.