Lesson 10 / 25
Checking for Bias with Slices
Compare quality across groups and languages instead of reporting one overall number.
Look at the slices
A system that is 92% accurate overall can be 97% accurate for one group and 70% for another. Report results by slice: language (English vs Hindi), dialect, region, age group, topic or any other attribute relevant to your users and permitted to be used for testing. Large gaps mean the system serves some people worse. Fix them with better data, prompts or evaluation sets, and decide what gap is acceptable before launch.
Beyond averages
An average score can hide a group that is badly served. Fairness and privacy need explicit checks and records.
Accuracy by slice
A tiny worked example. Overall accuracy is 0.75, but Hindi is at 0.5 while English is at 1.0. Always check counts too: tiny slices give noisy numbers.
results = [("en", True), ("en", True), ("en", True), ("en", True),
("hi", True), ("hi", False), ("hi", False), ("hi", True)]
print("overall", sum(ok for _, ok in results) / len(results))
for lang in ("en", "hi"):
sl = [ok for l, ok in results if l == lang]
print(lang, sum(sl) / len(sl), "n =", len(sl))
Output:
overall 0.75 en 1.0 n = 4 hi 0.5 n = 4
Include Hindi and mixed-language tests
Models are often tested mainly in English. If your users write Hindi, Hinglish or other Indian languages, build test cases in those languages too, or you will not see the quality gap until users do.
Quick check: Why report accuracy by slice instead of only overall?
- An overall average can hide a group that is served badly
- Slices are always more accurate
- Overall numbers are illegal
- It saves compute
Answer
An overall average can hide a group that is served badly — Gaps between groups only appear when results are broken down.