Lesson 16 / 25
LLM Judges and Human Agreement
Use a model to grade open-ended answers, but validate it against human labels with agreement statistics.
Judge the judge
Open-ended answers cannot be matched against a single expected string, so teams often use a second model as a judge with a clear rubric ("is the answer correct, complete and grounded in the sources?"). Judges can be biased toward longer answers or their own style, and they make mistakes. So compare the judge with human labels on a sample. Cohen's kappa measures agreement beyond chance: 0 means chance level, 1 means perfect agreement, and values around 0.4 to 0.6 are only moderate.
Cohen's kappa, run
I ran this on ten answers labelled by a human and by a judge. They agree on 8 of 10 (80%), but kappa is only about 0.52 once chance agreement is removed, so the judge is only moderately reliable.
def kappa(a, b):
n = len(a)
po = sum(x == y for x, y in zip(a, b)) / n
labels = set(a) | set(b)
pe = sum((a.count(l) / n) * (b.count(l) / n) for l in labels)
return round((po - pe) / (1 - pe), 3), round(po, 3)
human = ["good","good","bad","good","bad","good","good","bad","good","good"]
judge = ["good","good","bad","good","good","good","good","bad","good","bad"]
print(kappa(human, judge))
Output:
(0.524, 0.8)
Improve the rubric, then re-check
Vague rubrics give noisy judges. Write concrete criteria with examples of pass and fail, test the judge again against humans, and keep humans in the loop for the most important cases.
Quick check: What does Cohen's kappa measure?
- Cost per token
- Model speed
- Agreement between two raters beyond what chance would give
- Disk usage
Answer
Agreement between two raters beyond what chance would give — Kappa corrects raw agreement for the agreement expected by chance.