Lesson 14 / 25
Metrics: Accuracy, Precision, Recall, F1
Choose a metric that matches the cost of each kind of error.
Pick the metric that matches the risk
For yes/no decisions (is this message harmful? did the answer use a source?) count true positives (TP), false positives (FP) and false negatives (FN). Precision = TP / (TP + FP): of the items we flagged, how many were truly positive. Recall = TP / (TP + FN): of all truly positive items, how many we caught. F1 is their harmonic mean. Plain accuracy can mislead when one class is rare: a filter that never flags anything is 99% accurate if only 1% of messages are harmful.
Precision, recall and F1, run
I ran this with TP = 40, FP = 10, FN = 20: precision 0.8, recall about 0.667 and F1 about 0.727. Choose which matters more for your task: a safety filter usually favours recall, a spam filter often favours precision.
def prf(tp, fp, fn):
p = tp / (tp + fp)
r = tp / (tp + fn)
f = 2 * p * r / (p + r)
return round(p, 3), round(r, 3), round(f, 3)
print(prf(40, 10, 20))
Output:
(0.8, 0.667, 0.727)
Do not forget the base rate
If harmful messages are 1 in 100, a lazy rule that always says "safe" scores 99% accuracy and catches nothing. Always report precision and recall for the rare class, not only accuracy.
Quick check: What does recall measure?
- How many tokens were used
- Of flagged items, how many were truly positive
- How fast the model replies
- Of all truly positive items, how many were caught
Answer
Of all truly positive items, how many were caught — Recall is TP / (TP + FN): the share of real positives the system finds.