
Rubric drift: measuring guideline change mid-project
The guideline edit nobody logged is the most common root cause we find, and the easiest to prove once the rubric is versioned.
We think model performance is bounded by the data it learns from, and that the bound can be measured. These notes are what we measured, what changed because of it, and what each result does not establish.

Every layer on the site, with the published work it rests on and the limitation of that work, stated in the same place. Twenty-five references.
View the evidence
What the interview asks, why it never scores, and how a verified answer stops being asked. The same guide our experts read before they join a call.
Read the guide
The guideline edit nobody logged is the most common root cause we find, and the easiest to prove once the rubric is versioned.

An LLM judge that has not been checked against human adjudication is a confident random number generator.

A project-level coefficient is an average, and averages are exactly the wrong tool for finding the one person dragging a dataset down.
Schema and format checks that run before any human sees a record. We study what they reliably catch, which is malformed rows and exact duplicates, and what they cannot: a well-formed label that is simply wrong.
