Lesson 4 / 25
Moderation and Thresholds
Use a classifier score with a threshold and understand the precision-recall trade-off.
Where you draw the line
A moderation classifier gives a score (how likely the text is harmful). You choose a threshold above which you block or review. A low threshold catches almost all harmful text (high recall) but also blocks many harmless messages (low precision). A high threshold blocks only the clearest cases (high precision) but misses some harm (low recall). The right point depends on the cost of each mistake, so choose it with data, not by guessing.
Check before and after
Filters and checks guard both what goes into the model and what comes out of it.
Precision and recall at three thresholds, run
I ran this on eight labelled examples (1 = harmful). Raising the threshold from 0.5 to 0.85 raises precision from 0.8 to 1.0 but drops recall from 1.0 to 0.5.
scores = [(0.95,1),(0.9,1),(0.8,1),(0.7,0),(0.6,1),(0.4,0),(0.3,0),(0.2,0)] # (score, harmful?)
def at(th):
tp = sum(1 for s,y in scores if s>=th and y==1)
fp = sum(1 for s,y in scores if s>=th and y==0)
fn = sum(1 for s,y in scores if s<th and y==1)
return th, round(tp/(tp+fp),2) if tp+fp else None, round(tp/(tp+fn),2)
for th in (0.5, 0.65, 0.85):
print(at(th))
Output:
(0.5, 0.8, 1.0) (0.65, 0.75, 0.75) (0.85, 1.0, 0.5)
Use a review band
Instead of one cut-off, block above a high score, allow below a low one, and send the middle band to human review or a safer response. This spends human time only where the classifier is unsure.
Quick check: What happens to recall when you raise the moderation threshold?
- It always rises
- It usually falls, more harmful items are missed
- It becomes undefined
- It stays exactly the same
Answer
It usually falls, more harmful items are missed — A stricter cut-off blocks fewer items overall, so some real harm slips under it.