All posts
ResearchMar 11, 20261 minElenchLabs Research

Why Cohen's kappa hides your worst annotator

A project-level coefficient is an average, and averages are exactly the wrong tool for finding the one person dragging a dataset down.

An indexed dossier with a loupe on a highlighted line

The averaging problem

Mean pairwise kappa is computed over every annotator pair and then averaged. An annotator who agrees with nobody contributes to only a handful of those pairs, so their effect on the mean is diluted by exactly the thing that makes them a problem.

In a pool of twelve, one annotator at kappa 0.2 against everyone else moves the project mean by a few hundredths. The number still clears the bar. The labels are still bad.

What we report instead

Per-annotator kappa, sorted ascending, with the item count beside it. The count matters: kappa 0.31 over 400 items and kappa 0.31 over 9 items are different claims, and only one of them is worth acting on.

The correction

Weight by a measured reliability prior rather than treating the pool as exchangeable. An annotator whose interview already predicted drift on ambiguous items is not a surprise when they show up at the bottom of the table; they are a confirmed measurement.

What this does not establishPer-annotator kappa still says nothing about whether the pool as a whole read a clause the wrong way. That is the consensus layer's blind spot, and it needs a rubric fit review to find.