# Metrics: Accuracy, Precision, Recall, F1 — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/eval-metrics

> Choose a metric that matches the cost of each kind of error.

## Pick the metric that matches the risk

For yes/no decisions (is this message harmful? did the answer use a source?) count true positives (TP), false positives (FP) and false negatives (FN). **Precision** = TP / (TP + FP): of the items we flagged, how many were truly positive. **Recall** = TP / (TP + FN): of all truly positive items, how many we caught. **F1** is their harmonic mean. Plain **accuracy** can mislead when one class is rare: a filter that never flags anything is 99% accurate if only 1% of messages are harmful.

## Precision, recall and F1, run

I ran this with TP = 40, FP = 10, FN = 20: precision 0.8, recall about 0.667 and F1 about 0.727. Choose which matters more for your task: a safety filter usually favours recall, a spam filter often favours precision.

```python
def prf(tp, fp, fn):
    p = tp / (tp + fp)
    r = tp / (tp + fn)
    f = 2 * p * r / (p + r)
    return round(p, 3), round(r, 3), round(f, 3)

print(prf(40, 10, 20))
```

Output:

```
(0.8, 0.667, 0.727)
```

## Do not forget the base rate

If harmful messages are 1 in 100, a lazy rule that always says "safe" scores 99% accuracy and catches nothing. Always report precision and recall for the rare class, not only accuracy.

**Quiz:** What does recall measure?

- [ ] How many tokens were used
- [ ] Of flagged items, how many were truly positive
- [ ] How fast the model replies
- [x] Of all truly positive items, how many were caught

*Answer:* Of all truly positive items, how many were caught. Recall is TP / (TP + FN): the share of real positives the system finds.
