发表机构
University of California, Berkeley; National Bureau of Economic Research (NBER)(加州大学伯克利分校; 美国国家经济研究局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ORQA框架,通过连接O*NET职业与可信职业网站生成可溯源问答对,构建覆盖116个职业的基准,测试15个LLM发现职业间性能差异显著,为评估职业层面AI知识提供可扩展方法。
AI 中文摘要
我们提出了ORQA,一种用于测试大型语言模型中职业层面知识的方法。先前的方法要么通过任务定义将抽象的LLM技能映射到职业,要么利用专家知识,但专家知识难以大规模获取且成本高昂。ORQA通过将O*NET职业与可信的职业特定网站(如监管机构、许可机构、专业组织和政府出版物)连接起来,并将这些内容转化为可追溯来源的问答对,从而对这两种方法形成补充。自动化流程与人工审查相结合,产生了一套关于职业的高质量问题。通过我们的方法创建的问答集覆盖了SOC中所有21个主要职业组中的116个职业,包含来自187个不同网站的480个问题。每个问题旨在探究与该职业相关的现实世界技能问题。我们通过此方法测试了15个最先进的前沿和开放权重模型。Claude Opus 4.6、GPT-5.4和Claude Sonnet 4.6均表现最佳,准确率约为58-62%,而较小的开放权重模型则达到约33-41%的性能。不同职业间的性能差异显著。医疗相关职业达到最高性能(78%),而办公和行政支持职业约为40%。个别职业(如钣金工人和鱼类与野生动物管理员)的性能基本为零。我们还发现,开放式问题和按工资总额加权并不会显著影响模型在此基准上的排名。我们相信,利用现有的可信职业特定信息来测试LLM在专业领域的知识,可能是未来评估职业层面AI性能的一种可扩展且有用的方法。结果和数据可在以下网址获取:此http URL。
英文摘要
We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utilize expert knowledge which is difficult to obtain at scale and expensive. ORQA complements both of these methods by connecting O*NET occupations to trusted occupation-specific websites (such as regulatory agencies, licensing bodies, professional organizations, and government publications) and converting these into source-traceable question-answer pairs. A combination of an automated pipeline and human review produces a set of high quality questions about occupations. The question set created via our method covers 116 occupations from all 21 major groups in the SOC, with 480 questions sourced from 187 different websites. Each question is designed to probe a real-world skill question that is relevant to the occupation in question. We test 15 state-of-the-art frontier and open-weight models via this method. Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 all perform the best at approximately 58-62% while smaller open-weight models achieve approximately 33-41% performance. Performance varies significantly across occupations. Healthcare-related occupations achieve the highest performance (78%) while Office and Administrative Support achieve approximately 40%. Performance on individual occupations (e.g. Sheet Metal Workers and Fish and Game Wardens) is essentially zero. We also find that open-ended questions and weighting by wage bill do not significantly affect the ranking of models on this benchmark. We believe that leveraging existing trusted occupation-specific information to test LLM knowledge in professional domains may be a scalable and useful method for evaluating occupation-level AI performance in the future. Results and data are available at orqabench.org.
Comments45 pages, 17 figures, 6 tables. Data, code, and an interactive dashboard at orqabench.org