发表机构
AIRELab, Department of Educational Leadership and Policy Studies, University of Tennessee, Knoxville; Department of Computer Science and Engineering, BRAC University; Heinz College of Information Systems and Public Policy, Carnegie Mellon University(田纳西大学诺克斯维尔分校教育领导与政策研究系; BRAC大学计算机科学与工程系; 卡内基梅隆大学海因茨信息系统与公共政策学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文基于文化响应性评估框架,审计LLM基准SimpleQA和Chatbot Arena,发现结构性效度缺陷,并提出实用CR评估框架以改善评估公平性。
AI 中文摘要
LLM基准作为评估工具,影响着全球教育、劳动和公共服务领域的决策。本文借鉴Hood、Kirkhart和Hopson的文化响应性评估(CRE)框架,应用六维度CR评分标准,对OpenAI的SimpleQA(N=4,326个条目)和LMSYS Chatbot Arena(N=600个对话)进行了审计。每个SimpleQA问题都要求以英文档案验证作为其证据基础。一位评分者对哥伦比亚建国日期的关注占条目的2.70%,夸大了全球南方覆盖的表象。英文提示占Arena对话的76.3%,而国际电信联盟(ITU)估计全球互联网用户中英文使用者占比为25.9%。一个包含50个条目的反基准测试的平均CR缺陷得分比SimpleQA低近三倍(Cohen's d=1.01)。本文提出了一个实用的CR评估框架。这些是结构性效度失败,而非偶然的测量问题,对那些知识传统未被这些工具设计来识别的社区产生了直接影响。
英文摘要
LLM benchmarks function as evaluation instruments, informing decisions that affect education, labor, and public services worldwide. Drawing on Hood, Kirkhart, and Hopson's culturally responsive evaluation (CRE) frameworks, this paper applies a six-dimension CR rubric to audit OpenAI's SimpleQA (N = 4,326 items) and the LMSYS Chatbot Arena (N = 600 conversations). Every SimpleQA question requires English-language archival verification as its evidentiary basis. A single rater's preoccupation with Colombian founding dates accounts for 2.70% of items, inflating the appearance of Global South coverage. English-language prompts constitute 76.3% of Arena conversations, against an International Telecommunication Union (ITU)-estimated 25.9% share of global internet users. A 50-item counter-benchmark scored a mean CR deficit nearly three times lower than SimpleQA (Cohen's d = 1.01). The paper proposes a practical CR evaluation framework. These are structural validity failures, not incidental measurement problems, with direct consequences for communities whose knowledge traditions these instruments were not built to see.
Comments36 pages, 8 figures