Research

The published work each quality layer rests on, and what that work does not establish, stated together in the same place.

Quality layer

Automated checks

Even the benchmark datasets the field treats as ground truth carry percent-level label errors, and in practitioners' own accounts the failures that follow are opaque and delayed.

  1. [1]

    At least 3.3% of labels are wrong on average across ten standard benchmark test sets, and at least 6% of the ImageNet validation set. The errors invert model selection: on corrected labels ResNet-18 outperforms ResNet-50 once the share of originally mislabelled test examples rises by 6%.

  2. [2]

    In a production data-validation deployment across more than 700 pipelines, an unexpected new feature column fired on 10% of pipelines and a missing one on 6%. Schema-driven model unit tests ran over 80,000 times in a month and 6% of runs failed, each failure a wrong assumption about the data.

  3. [3]

    Of 53 practitioners building high-stakes AI, 92% had hit at least one data cascade and 45.3% had hit two or more inside a single project, failures the authors characterise as opaque, delayed, and largely avoidable.

Quality layer

Golden set

Seeding known-answer items turns an unverifiable pile of labels into a measurable one, and without it, a startling share of submitted work is not a genuine attempt at the task.

  1. [4]

    Adding four verifiable questions ahead of an otherwise identical rating task cut invalid responses from 48.6% to 2.5%, raised median time on task from 1:30 to 4:06, and lifted agreement with expert raters from r = 0.50 to r = 0.66.

  2. [5]

    On a 500-page labelling task with 339 workers, removing workers by a cost-aware score cut annotation cost by 30% while raising quality from 0.95 to 0.998; spammers had produced 30% of submitted answers. The same filtering without the cost-aware score cut costs by 1% and left quality effectively unchanged.

  3. [6]

    Two embedded screening questions identified 764 of 1,962 participants (38.9%) as not answering conscientiously, on a task paying $4 for thirty minutes.

  4. [7]

    Across five experiments with 255 workers, between 7% and 27% failed the gold training screen before reaching production data.

Quality layer

Consensus

A handful of independent annotators is a workable substitute for an expert (about four non-expert annotations reached single-expert agreement across seven tasks), and modelling whom to trust beats counting votes.

  1. [8]

    Pooled across seven tasks, an average of four non-expert annotations per example matched the inter-annotator agreement of one expert. On word sense disambiguation the crowd's only disagreement with gold turned out to be an error in the gold standard, outvoted 9 to 1.

  2. [9]

    Majority vote improves on a single label only while the average annotator is better than chance. At 70% per-annotator accuracy, moving from one labeller to three lifts integrated quality by about 0.1; at 90%, moving from three to eleven buys almost nothing.

  3. [10]

    Modelling annotator competence rather than counting votes raised accuracy on the same data from 0.90 to 0.93 on recognising textual entailment, and to 0.98 when a quarter of items were deferred instead of forced.

  4. [11]

    The original method, fitted to five anaesthetists rating 45 patients with no ground truth available at all, recovered per-observer error rates that differ sharply: one observer recorded a true category-4 patient correctly only 44% of the time.

Quality layer

AI review

A strong judge, checked against human decisions, agrees with people about as often as two people agree with each other, and reverses itself when nobody is checking.

  1. [13]

    GPT-4 matched human preference on 85% of MT-Bench comparisons and 87% on Chatbot Arena, against 81% agreement between the human raters themselves.

  2. [13]

    The same judge held its verdict after the two answers were swapped only 65.0% of the time, and scored its own answers about 10 points above the human win rate.

  3. [14]

    In a separate setup with a different judge and candidate pool, changing the order of the candidates alone flipped the outcome on 66 of 80 queries; the two studies bound the effect rather than agreeing on its size.

  4. [15]

    The best prompted GPT-4 evaluator reached Spearman 0.514 against human summarization judgments: the state of the art, and still a moderate correlation.

Quality layer

Rubric fit review

Agreement is a property of the instrument. How the scheme is written and how annotators are trained move it measurably, and both are things you control.

  1. [17]

    Across 96 studies and 346 data points, trained annotators averaged 81% agreement against 70% untrained, and 86% under intensive training. Agreement fell reliably as the scheme grew categories (β = −0.28, p < 0.001), replicated in all three domains.

  2. [18]

    Revising the guideline document alone raised inter-annotator agreement on WNUT-17 from 0.593 to 0.84.

  3. [19]

    Low agreement on subjective tasks often marks a guideline that never chose between a descriptive standard and a prescriptive one.

Quality layer

Adversarial testing

Held-out accuracy overstates capability. People probing on purpose find failures that a benchmark score hides.

  1. [20]

    On three commercial sentiment APIs, negation at the end of a sentence failed 100% of the time for two of them and 90.4% for the third. Practitioners given the method wrote twice as many tests and found nearly three times as many bugs.

  2. [21]

    Small meaning-changing edits by experts, which the authors stress were not adversarial, dropped performance by up to 25% across 10 datasets.

  3. [22]

    RoBERTa scoring 92.6% on SNLI fell to 48.9% and 44.4% on the second and third rounds of adversarially collected data.

  4. [23]

    Across 38,961 red-team attacks, only RLHF-trained models became harder to attack as they scaled; the other three model types were flat.

References

  1. [1]

    Northcutt, Athalye, Mueller (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. NeurIPS 2021 Datasets and Benchmarks.

  2. [2]

    Breck, Polyzotis, Roy, Whang, Zinkevich (2019). Data Validation for Machine Learning. MLSys 2019.

  3. [3]

    Sambasivan, Kapania, Highfill, Akrong, Paritosh, Aroyo (2021). Everyone wants to do the model work, not the data work: Data Cascades in High-Stakes AI. CHI 2021.

  4. [4]

    Kittur, Chi, Suh (2008). Crowdsourcing User Studies with Mechanical Turk. CHI 2008.

  5. [5]

    Ipeirotis, Provost, Wang (2010). Quality Management on Amazon Mechanical Turk. HCOMP 2010.

  6. [6]

    Downs, Holbrook, Sheng, Cranor (2010). Are Your Participants Gaming the System? Screening Mechanical Turk Workers. CHI 2010.

  7. [7]

    Le, Edmonds, Hester, Biewald (2010). Ensuring Quality in Crowdsourced Search Relevance Evaluation: The Effects of Training Question Distribution. SIGIR 2010 Workshop on Crowdsourcing for Search Evaluation.

  8. [8]

    Snow, O'Connor, Jurafsky, Ng (2008). Cheap and Fast, But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. EMNLP 2008.

  9. [9]

    Sheng, Provost, Ipeirotis (2008). Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers. KDD 2008.

  10. [10]

    Hovy, Berg-Kirkpatrick, Vaswani, Hovy (2013). Learning Whom to Trust with MACE. NAACL-HLT 2013.

  11. [11]

    Dawid, Skene (1979). Maximum Likelihood Estimation of Observer Error-rates using the EM Algorithm. Journal of the Royal Statistical Society Series C, 28(1).

  12. [12]

    Plank (2022). The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. EMNLP 2022.

  13. [13]

    Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks.

  14. [14]

    Wang, Li, Chen, Cai, Zhu, Lin, Cao, Liu, Liu, Sui (2023). Large Language Models are not Fair Evaluators. ACL 2024.

  15. [15]

    Liu, Iter, Xu, Wang, Xu, Zhu (2023). G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP 2023.

  16. [16]

    Thakur, Choudhary, Ramayapally, Vaidyanathan, Hupkes (2024). Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. arXiv.

  17. [17]

    Bayerl, Paul (2011). What Determines Inter-Coder Agreement in Manual Annotations? A Meta-Analytic Investigation. Computational Linguistics 37(4).

  18. [18]

    Bibal, Gerlek, Muric, Boschee, Fincke, Ross, Minton (2025). Automating Annotation Guideline Improvements using LLMs: A Case Study. CoMeDi workshop, ACL Anthology.

  19. [19]

    Röttger, Vidgen, Hovy, Pierrehumbert (2022). Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. NAACL 2022.

  20. [20]

    Ribeiro, Wu, Guestrin, Singh (2020). Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. ACL 2020.

  21. [21]

    Gardner et al. (2020). Evaluating Models' Local Decision Boundaries via Contrast Sets. Findings of EMNLP 2020.

  22. [22]

    Nie, Williams, Dinan, Bansal, Weston, Kiela (2020). Adversarial NLI: A New Benchmark for Natural Language Understanding. ACL 2020.

  23. [23]

    Ganguli, Lovitt, Kernion, Askell, Bai et al. (2022). Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv.

  24. [24]

    Kaushik, Kiela, Lipton, Yih (2021). On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized Study. ACL-IJCNLP 2021.

  25. [25]

    Wallace, Williams, Jia, Kiela (2022). Analyzing Dynamic Adversarial Training Data in the Limit. Findings of ACL 2022.

Numbered by first appearance above. Every reference links to the paper it names; a finding is quoted at the precision the paper reports it.