- [1]
Northcutt, Athalye, Mueller (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. NeurIPS 2021 Datasets and Benchmarks.
- [2]
Breck, Polyzotis, Roy, Whang, Zinkevich (2019). Data Validation for Machine Learning. MLSys 2019.
- [3]
Sambasivan, Kapania, Highfill, Akrong, Paritosh, Aroyo (2021). Everyone wants to do the model work, not the data work: Data Cascades in High-Stakes AI. CHI 2021.
- [4]
Kittur, Chi, Suh (2008). Crowdsourcing User Studies with Mechanical Turk. CHI 2008.
- [5]
Ipeirotis, Provost, Wang (2010). Quality Management on Amazon Mechanical Turk. HCOMP 2010.
- [6]
Downs, Holbrook, Sheng, Cranor (2010). Are Your Participants Gaming the System? Screening Mechanical Turk Workers. CHI 2010.
- [7]
Le, Edmonds, Hester, Biewald (2010). Ensuring Quality in Crowdsourced Search Relevance Evaluation: The Effects of Training Question Distribution. SIGIR 2010 Workshop on Crowdsourcing for Search Evaluation.
- [8]
Snow, O'Connor, Jurafsky, Ng (2008). Cheap and Fast, But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. EMNLP 2008.
- [9]
Sheng, Provost, Ipeirotis (2008). Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers. KDD 2008.
- [10]
Hovy, Berg-Kirkpatrick, Vaswani, Hovy (2013). Learning Whom to Trust with MACE. NAACL-HLT 2013.
- [11]
Dawid, Skene (1979). Maximum Likelihood Estimation of Observer Error-rates using the EM Algorithm. Journal of the Royal Statistical Society Series C, 28(1).
- [12]
Plank (2022). The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. EMNLP 2022.
- [13]
Zheng, Chiang, Sheng, Zhuang, Wu, Zhuang, Lin, Li, Li, Xing, Zhang, Gonzalez, Stoica (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks.
- [14]
Wang, Li, Chen, Cai, Zhu, Lin, Cao, Liu, Liu, Sui (2023). Large Language Models are not Fair Evaluators. ACL 2024.
- [15]
Liu, Iter, Xu, Wang, Xu, Zhu (2023). G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP 2023.
- [16]
Thakur, Choudhary, Ramayapally, Vaidyanathan, Hupkes (2024). Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. arXiv.
- [17]
Bayerl, Paul (2011). What Determines Inter-Coder Agreement in Manual Annotations? A Meta-Analytic Investigation. Computational Linguistics 37(4).
- [18]
Bibal, Gerlek, Muric, Boschee, Fincke, Ross, Minton (2025). Automating Annotation Guideline Improvements using LLMs: A Case Study. CoMeDi workshop, ACL Anthology.
- [19]
Röttger, Vidgen, Hovy, Pierrehumbert (2022). Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. NAACL 2022.
- [20]
Ribeiro, Wu, Guestrin, Singh (2020). Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. ACL 2020.
- [21]
Gardner et al. (2020). Evaluating Models' Local Decision Boundaries via Contrast Sets. Findings of EMNLP 2020.
- [22]
Nie, Williams, Dinan, Bansal, Weston, Kiela (2020). Adversarial NLI: A New Benchmark for Natural Language Understanding. ACL 2020.
- [23]
Ganguli, Lovitt, Kernion, Askell, Bai et al. (2022). Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv.
- [24]
Kaushik, Kiela, Lipton, Yih (2021). On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized Study. ACL-IJCNLP 2021.
- [25]
Wallace, Williams, Jia, Kiela (2022). Analyzing Dynamic Adversarial Training Data in the Limit. Findings of ACL 2022.
Numbered by first appearance above. Every reference links to the paper it names; a finding is quoted at the precision the paper reports it.