Lesson 8 / 29

LLM-as-Judge: Useful, Biased, Must Be Calibrated

Use a model to grade open-ended output, and check the grader.

A judge needs its own evaluation

For open-ended output (summaries, explanations, chat replies) exact matching fails, so teams use a strong model as a judge with a rubric ("rate faithfulness 1 to 5; list unsupported claims"). It scales well but has known weaknesses: position bias (favouring the first answer shown), length bias (favouring longer answers), self-preference (favouring text from its own model family), and inconsistency. Treat the judge as a measuring instrument that must be calibrated: label 50 to 100 cases by humans, measure agreement using a chance-corrected statistic such as Cohen's kappa (raw agreement can mislead: a judge that always says "good" agrees 60% of the time here but has kappa 0), randomise or swap answer order, give clear rubrics with examples, prefer pairwise or checklist-style judgments to vague overall scores, and re-check the judge when you change it. Keep humans in the loop for high-stakes decisions.

Agreement versus kappa, run

I ran this with plain Python 3 (standard library only). A judge agrees with the human on 8 of 10 cases (kappa 0.58, moderate). A lazy judge that always answers "good" agrees on 6 of 10 just because 60% of the human labels are "good", but its kappa is 0.00: it adds no information.

def cohen_kappa(a, b):
    n = len(a); po = sum(x == y for x, y in zip(a, b)) / n
    labels = set(a) | set(b)
    pe = sum((a.count(l) / n) * (b.count(l) / n) for l in labels)
    return (po - pe) / (1 - pe) if pe < 1 else 1.0

human = ["good", "bad", "good", "good", "bad", "bad", "good", "bad", "good", "good"]
judge = ["good", "bad", "good", "bad",  "bad", "bad", "good", "good", "good", "good"]
agree = sum(h == j for h, j in zip(human, judge))
print(f"raw agreement: {agree}/10 | Cohen's kappa: {cohen_kappa(human, judge):.2f}")
always_good = ["good"] * 10
print(f"a judge that always says 'good': agreement {sum(h == 'good' for h in human)}/10, kappa {cohen_kappa(human, always_good):.2f}")

Output:

raw agreement: 8/10 | Cohen's kappa: 0.58
a judge that always says 'good': agreement 6/10, kappa 0.00

Position bias and the order-swap fix (simulation), run

I ran this with plain Python 3 (standard library only). This is a simulation with invented numbers and seeded random draws, so it shows the shape of the effect, not measurements of any product. Two answers are exactly equally good, but the simulated judge prefers whichever is shown first 70% of the time. A looks like it wins about 85% of the time when shown first and 15% when shown second. Averaging over both orders removes the bias and recovers 0.50, which is the true value.

# A and B are EXACTLY equally good. A judge with position bias picks whichever answer is shown FIRST 70% of the time
# (simulated); otherwise it flips a coin.
import random
rng = random.Random(1)
def biased_judge():
    return "first" if rng.random() < 0.7 else rng.choice(["first", "second"])

N = 20000
a_first  = sum(biased_judge() == "first"  for _ in range(N)) / N       # A shown first
a_second = sum(biased_judge() == "second" for _ in range(N)) / N       # A shown second
print("A wins when shown first :", round(a_first, 3))
print("A wins when shown second:", round(a_second, 3))
print("average over both orders:", round((a_first + a_second) / 2, 3), "(true value is 0.5)")

Output:

A wins when shown first : 0.852
A wins when shown second: 0.152
average over both orders: 0.502 (true value is 0.5)

Give the judge a rubric with examples

A clear scoring guide with graded examples makes judgments more consistent.

Quick check: What is a simple way to reduce position bias in pairwise judging?

  • Use a shorter rubric
  • Always show the better answer first
  • Judge each pair in both orders and combine the results
  • Ignore the judge
Answer

Judge each pair in both orders and combine the results — Swapping order cancels a bias that favours position.