大语言模型能否预测个体对社会政策的行为响应?以中国灵活就业人员的养老金参保预测为例
Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers
浏览论文内容
中文总结 AI 辅助
本研究提出领域专用大语言模型FlexPension-LLM,结合DKI-RDistill方法,在灵活就业人员养老金参保预测任务中实现优异性能,为社会政策影响评估提供新工具。
中文摘要 AI 辅助
评估社会政策变动的影响是政策制定者公认的难题,计量经济学方法在推演假设场景时可能不可靠,而实地试点项目成本极高。本文提出将大语言模型(LLMs)用作政策评估工具,该工具改编自通用模型。我们推出FlexPension-LLM,这是首个针对中国灵活就业人员分层养老金参保预测任务的领域专用大语言模型,并引入DKI-RDistill,该方法将基于政策的提示线索注入提示词,包括Probit推导的边际效应和户籍-省份养老金规则。该方法随后使用LoRA/SFT将基于理由增强的监督提炼为开放权重的MoE学生模型,同时通过在真实标签下重新生成案例来修正教师模型的错误。在CHFS 2019的盲分数据集上,FlexPension-LLM取得0.9316的综合F1值,超过其Claude Sonnet 4.5教师模型和17个基准模型中的15个,且与Claude Opus 4.6在统计上无差异。在四个外部调查中,它的综合F1值平均为0.7549,且在最强系统中表现范围最窄。组件分析显示,性能提升主要来自基于政策的线索注入和经错误过滤的监督,而理由则提供了可对照政策规则核查的决策轨迹。
英文摘要
Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highly costly. In this paper, we propose using large language models (LLMs) as policy-assessment tools adapted from general-purpose models. We present FlexPension-LLM, the first domain-specialized large language model for a hierarchical pension-enrollment prediction task among flexible workers in China, and introduce DKI-RDistill, which injects policy-grounded cues into the prompt, including Probit-derived marginal effects and hukou-province pension rules. The method then uses LoRA/SFT to distill rationale-augmented supervision into an open-weight MoE student, with teacher errors corrected by regenerating those cases under ground-truth labels. On a CHFS 2019 blind split, FlexPension-LLM achieves 0.9316 Composite F1, surpassing its Claude Sonnet 4.5 teacher and 15 of 17 baselines, and is statistically indistinguishable from Claude Opus 4.6. Across four external surveys, it averages 0.7549 Composite F1 and shows the narrowest performance range among the strongest systems. Component analysis shows that gains come mainly from policy-grounded cue injection and error-filtered supervision, while rationales provide decision traces that can be checked against policy rules.
发表机构
- School of Economics and Management, Tsinghua University(清华大学经济管理学院)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- School of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。