Why Cohen's kappa hides your worst annotator
A project-level coefficient is an average, and averages are exactly the wrong tool for finding the one person dragging a dataset down.

The averaging problem
Mean pairwise kappa is computed over every annotator pair and then averaged. An annotator who agrees with nobody contributes to only a handful of those pairs, so their effect on the mean is diluted by exactly the thing that makes them a problem.
In a pool of twelve, one annotator at kappa 0.2 against everyone else moves the project mean by a few hundredths. The number still clears the bar. The labels are still bad.
What we report instead
Per-annotator kappa, sorted ascending, with the item count beside it. The count matters: kappa 0.31 over 400 items and kappa 0.31 over 9 items are different claims, and only one of them is worth acting on.
The correction
Weight by a measured reliability prior rather than treating the pool as exchangeable. An annotator whose interview already predicted drift on ambiguous items is not a surprise when they show up at the bottom of the table; they are a confirmed measurement.