Lesson 11 / 26

Measuring Fairness

Compute selection-rate ratios and true-positive-rate gaps between groups, and know their limits.

Two simple lenses

Disparate impact ratio: divide the favourable-outcome rate of one group by that of the group with the highest rate. A ratio below about 0.8 (the "four-fifths rule of thumb" used in some US employment guidance) is a warning sign, not a legal verdict everywhere. Equal opportunity compares the true positive rate across groups: among people who truly qualify, does each group get approved equally often? Different fairness measures can conflict, and none replaces judgement about context, so choose metrics with domain and legal experts and investigate causes, not only numbers.

Disparate impact and TPR gap, run

I ran this with made-up numbers. Group A is selected 60% of the time and group B 36%, a ratio of 0.6, which flags a problem. Among qualified people, 90% of group A but only 60% of group B are approved, a 30-point gap.

a = 60 / 100          # group A selection rate
b = 36 / 100          # group B selection rate
ratio = round(b / a, 2)
print(a, b, ratio, "FLAG" if ratio < 0.8 else "ok")

def tpr(tp, fn): return tp / (tp + fn)
g1, g2 = tpr(45, 5), tpr(30, 20)   # qualified people approved / rejected
print(round(g1, 2), round(g2, 2), round(g1 - g2, 2))

Output:

0.6 0.36 0.6 FLAG
0.9 0.6 0.3

A flag starts an investigation

A bad ratio does not tell you why. Check the data (is one group under-represented or labelled differently?), the features used, and the process around the model before deciding on a fix.

Quick check: Group A is selected 60% of the time and group B 36%. What is the disparate impact ratio?

  • 96
  • 1.67
  • 0.24
  • 0.6
Answer

0.6 — 36 divided by 60 is 0.6, below the 0.8 rule of thumb, so it warrants investigation.