All posts
MethodApr 2, 20261 minElenchLabs Research

Calibrating an LLM judge

An LLM judge that has not been checked against human adjudication is a confident random number generator.

Polished cross-section of layered green and cream stone

The setup

A judge is scored blind against a set of items humans have already adjudicated to consensus, mixed into a larger batch so it cannot condition on being tested. The adjudications are the reference; the judge's agreement with them is the calibration.

What moves the number

Rubric specificity, overwhelmingly. A judge reading a one-paragraph rubric and the same judge reading one that enumerates its edge cases will agree with the reference by different margins. Model choice matters less than the document it is reading.

The ambiguous tail matters too: items the human adjudicators themselves needed a tiebreak on are where a judge diverges most. Disagreement concentrates in the same places for both.

Why we re-run it every batch

Calibration is not a property of the model. It is a property of the model, the rubric version, and the data distribution together, so it is measured per batch and recorded with all three.

What this does not establishA calibrated judge is calibrated for one rubric version and one data distribution. Change either and the number is stale, which is why it is re-measured per batch rather than once at setup.