Calibrating an LLM judge
An LLM judge that has not been checked against human adjudication is a confident random number generator.

The setup
A judge is scored blind against a set of items humans have already adjudicated to consensus, mixed into a larger batch so it cannot condition on being tested. The adjudications are the reference; the judge's agreement with them is the calibration.
What moves the number
Rubric specificity, overwhelmingly. A judge reading a one-paragraph rubric and the same judge reading one that enumerates its edge cases will agree with the reference by different margins. Model choice matters less than the document it is reading.
The ambiguous tail matters too: items the human adjudicators themselves needed a tiebreak on are where a judge diverges most. Disagreement concentrates in the same places for both.
Why we re-run it every batch
Calibration is not a property of the model. It is a property of the model, the rubric version, and the data distribution together, so it is measured per batch and recorded with all three.