arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19375cs.CYcs.AI

语言模型的经济评估

Economic Evaluations of Language Models

Alexander Wan, Stephane Hatgis-Kessell, Tomás Aguirre, Percy Liang, Rishi Bommasani

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型在经济价值任务上的表现,通过EconEvals开源套件评估,结合真实用户查询与合成数据,成本低且覆盖度好,还引入暴露度量估计时间节省情况,发现当前模型有潜力但存在使用滞后及隐私等瓶颈,引入基础设施推断其对劳动力市场的影响。

中文摘要 AI 辅助

语言模型能完成具有经济价值的工作,但目前尚未全面评估其在各项经济价值任务上的表现。我们引入EconEvals这一开源评估套件,通过真实用户查询及合成数据来衡量与美国劳动力经济中任务、工作活动和职业相关的能力。我们的评估在成本降低500倍的情况下,比OpenAI的GDPval基准有更好的覆盖度。还引入基于模拟的暴露度量来估计语言模型能力能为美国所有职业任务节省的时间。结果表明当前模型能为部分职业节省大量时间,但存在使用滞后情况,且隐私和专有系统是限制进一步节省时间的瓶颈。我们引入了可适应的基础设施来推断语言模型对劳动力市场的影响。

英文摘要

Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.

发表机构

  • Stanford University(斯坦福大学)
  • University of São Paulo(圣保罗大学)

机构由 AI 辅助整理,请以论文原文为准。

↑