发表机构
University of Michigan; Cornell University(密歇根大学; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文借鉴陈述偏好经济学的有效性框架,提出一种无需真实标签评估语言模型的方法,并通过水质调查数据对六个模型进行测试,区分了模型的理论有效性与收敛效度。
AI 中文摘要
现在向大型语言模型提出的许多问题没有正确答案可供评分:例如政策的价值、用户应选择的选项、如何权衡相互竞争的价值观。陈述偏好经济学几十年来一直面临这个问题。它在不知道真实价值的情况下,通过有效性及相关概念(内容效度、构念效度、标准效度、信度、激励相容性和后果性)的框架来评判调查回答。我们认为这一框架是评估语言模型的通用方法,并阐述了每个概念对LLM评估的意义。我们使用一项已发表的水质陈述偏好经济评估调查(Vossler et al. 2023)对六个模型进行了测试,以展示该方法。在这个经济应用中,有效性测试采用经济理论预测的形式:需求应向下倾斜,支付意愿应对商品范围和收入做出响应。这些测试将模型明显区分开来。两个较旧的模型在家庭收入水平为75,000美元时未能通过最基本的测试,而最新的两个模型通过了我们能够评分的所有理论有效性测试,但在收敛效度上存在分歧。通过有效性测试表明模型的回答是连贯的,而非正确的。
英文摘要
Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of \$75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model's answers are coherent, not that they are correct.