Lesson 14 / 25

Metrics: Accuracy, Precision, Recall, F1

Choose a metric that matches the cost of each kind of error.

Pick the metric that matches the risk

For yes/no decisions (is this message harmful? did the answer use a source?) count true positives (TP), false positives (FP) and false negatives (FN). Precision = TP / (TP + FP): of the items we flagged, how many were truly positive. Recall = TP / (TP + FN): of all truly positive items, how many we caught. F1 is their harmonic mean. Plain accuracy can mislead when one class is rare: a filter that never flags anything is 99% accurate if only 1% of messages are harmful.

Precision, recall and F1, run

I ran this with TP = 40, FP = 10, FN = 20: precision 0.8, recall about 0.667 and F1 about 0.727. Choose which matters more for your task: a safety filter usually favours recall, a spam filter often favours precision.

def prf(tp, fp, fn):
    p = tp / (tp + fp)
    r = tp / (tp + fn)
    f = 2 * p * r / (p + r)
    return round(p, 3), round(r, 3), round(f, 3)

print(prf(40, 10, 20))

Output:

(0.8, 0.667, 0.727)

Do not forget the base rate

If harmful messages are 1 in 100, a lazy rule that always says "safe" scores 99% accuracy and catches nothing. Always report precision and recall for the rare class, not only accuracy.

Quick check: What does recall measure?

  • How many tokens were used
  • Of flagged items, how many were truly positive
  • How fast the model replies
  • Of all truly positive items, how many were caught
Answer

Of all truly positive items, how many were caught — Recall is TP / (TP + FN): the share of real positives the system finds.