Research

Data quality is a measurement, not a promise

We think model performance is bounded by the data it learns from, and that the bound can be measured. These notes are what we measured, what changed because of it, and what each result does not establish.

All posts, newest first
Filter by kind

Calibrating an LLM judge

ElenchLabsMethodApr 2, 2026

An LLM judge that has not been checked against human adjudication is a confident random number generator.

Read post
6

Quality layers we research

Automated checks

Schema and format checks that run before any human sees a record. We study what they reliably catch, which is malformed rows and exact duplicates, and what they cannot: a well-formed label that is simply wrong.

Automated checks