# Refusal Calibration — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/io-refusal-calibration

> Measure both under-refusal of harmful requests and over-refusal of harmless ones.

## Two ways to be wrong

A safe assistant must refuse some requests, but refusing too much is also a failure. **Under-refusal** means helping with something harmful. **Over-refusal** means declining a harmless request (for example a nurse asking about medication doses, or a student asking how phishing works to defend against it). Track both rates on a test set that contains harmful and benign examples, and tune instructions and filters until both are acceptable.

## Two rates, run

I ran this on eight cases (3 harmful, 5 benign; the boolean says whether the assistant refused). One of three harmful requests was answered (under-refusal 0.33) and one of five benign requests was refused (over-refusal 0.2).

```python
cases = [("harmful",True),("harmful",True),("harmful",False),
         ("benign",True),("benign",False),("benign",False),("benign",False),("benign",False)]

harmful = [ref for kind, ref in cases if kind == "harmful"]
benign  = [ref for kind, ref in cases if kind == "benign"]
under = sum(1 for r in harmful if not r) / len(harmful)
over  = sum(1 for r in benign if r) / len(benign)
print(round(under, 2), round(over, 2))
```

Output:

```
0.33 0.2
```

## Refuse helpfully

A good refusal is brief, non-judgemental and offers a safe alternative or a resource, rather than a lecture. Poor refusals push users to less careful tools.

**Quiz:** What is over-refusal?

- [x] Declining a harmless request
- [ ] Answering a harmful request
- [ ] Refusing to start
- [ ] Running out of tokens

*Answer:* Declining a harmless request. Over-refusal makes the product less useful and is measured alongside under-refusal.
