arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HouseholdBench:评估大型语言模型作为家庭经济行为预测器的性能

HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior

Jin Huang, Diego Ferreras Garrucho, Yutong Xie, Walter M. Yuan, Qiaozhu Mei, Chen Lian, Jonathon Hazell

arXiv 2610.07563首次发表:更新:

发表机构

University of Michigan, Ann Arbor; London School of Economics; MobLab Inc; University of California, Berkeley(密歇根大学安娜堡分校; 伦敦政治经济学院; MobLab公司; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HouseholdBench评估,整合6项美国调查和32项预测任务,测试LLM预测家庭经济行为及政策响应的能力,发现多数LLM优于无变化基线,微调和聚合预测可提升开源模型性能。

AI 中文摘要

大型语言模型(LLMs)有潜力实现经济学中的一个关键目标:在各种情境下建立家庭决策的定量模型。然而,现有的评估仅涵盖少数调查和结果,并且未研究家庭如何适应不断变化的经济条件。我们引入了一项新的评估——HouseholdBench,它整合了6项美国住户调查和32项预测任务,涵盖与消费、收入、劳动、预期和住房相关的数值型、类别型和概率型结果。利用过去的行为、人口统计特征和宏观经济条件,这些任务测试LLMs是否能预测行为,包括家庭如何适应各种政策的变化。我们评估了13个专有和开源权重的LLMs,并与一个无变化基线和梯度提升树模型进行比较。大多数LLMs优于无变化基线,包括在政策响应任务中——最佳模型将数值型结果的误差降低了12.2%。在大多数任务中,梯度提升树排名第一;领先的专有LLMs接近其性能,但开源权重模型则落后。LLMs在不同任务中表现出系统性的过度预测和不足预测。我们确定了使一个40亿参数的开源权重模型能够匹配专有模型性能的方法:微调和对每个观测值聚合16个预测。这些改进也泛化到政策响应任务,而这些任务被排除在微调之外。我们在我们的网站上发布了数据集、代码和排行榜:此https URL。

英文摘要

Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6 U.S. household surveys and 32 prediction tasks spanning numeric, categorical and probabilistic outcomes, related to consumption, income, labor, expectations, and housing. Using past behavior, demographics and macroeconomic conditions, the tasks test whether LLMs predict behavior, including how households adjust to changes in various policies. We evaluate 13 proprietary and open-weight LLMs against a no-change baseline and a gradient-boosted tree model. Most LLMs outperform the no-change baseline, including for policy response tasks -- with the best model lowering error for numeric outcomes by 12.2%. Across most tasks, gradient-boosted trees rank first; leading proprietary LLMs approach their performance, but open-weight models lag. LLMs exhibit systematic over- and underprediction across different tasks. We identify methods that enable a 4 billion parameter open-weight model to match proprietary models' performance: fine-tuning and aggregating 16 predictions per observation. Improvements generalize to policy-response tasks, which are excluded from fine-tuning. We release our datasets, code, and leaderboard on our website: https://jn-huang.github.io/householdbench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑