arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向高效且高性价比LLM社会调查模拟的自适应资源分配

Adaptive Resource Allocation for Effective and Efficient LLM Social Survey Simulation

Yuanzi Li, Xueyang Feng, Junhao Wang, Lei Wang, Xu Chen

arXiv 2609.35216首次发表:更新:

发表机构

Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有LLM社会调查模拟统一分配资源的低效问题,提出自适应框架E2Sim,联合选择模型与历史预算,经实验验证可同时提升准确率并降低成本。

AI 中文摘要

大语言模型(LLM)让可扩展的社会调查模拟成为可能,但现有流程通常对每个受访者-问题请求都使用同一款性能强劲的通用模型,以及固定且往往篇幅较长的受访者历史记录。这种统一方法忽略了三个因素:首先,更强的模型可能依赖自身知识而非受访者特定证据,同时成本更高;其次,当证据有限时补充历史记录会有帮助,但无关回复可能引入噪声并增加输入长度;第三,最优模型与历史记录预算可能相互影响。我们提出E2Sim,这是一种自适应资源分配框架,可为每个请求联合选择模型和历史记录预算。给定受访者人设、排序后的回复历史以及目标问题,一个轻量策略会预测每种配置的准确率和成本,并选择最合适的配置。我们采用受访者历史丢弃-交换增强来提升对不完整或可变历史的鲁棒性,还采用基于裕度的课程学习,从较易的分配决策逐步过渡到较难的决策。在四个真实世界社会调查数据集、多个模型池以及不同历史预算空间上的实验表明,该方法优于最优固定配置,准确率最高提升6.3个百分点,成本最高降低55.3%。代码可在https://anonymous.4open.science/r/E2Sim-DE66获取。

英文摘要

Large Language Models (LLMs) enable scalable social survey simulation, yet existing pipelines typically use the same strong general-purpose model and a fixed, often large, respondent history for every respondent-question request. This uniform approach overlooks three factors. First, stronger models may rely on their own knowledge rather than respondent-specific evidence while costing more. Second, additional history can help when evidence is limited, but irrelevant responses may add noise and increase input length. Third, the preferred model and history budget can depend on each other. We propose E2Sim, an adaptive resource allocation framework that jointly selects a model and history budget for each request. Given a respondent persona, ranked response history, and target question, a lightweight policy predicts the accuracy and cost of each configuration and selects the most suitable one. We use respondent-history drop-and-swap augmentation to improve robustness to incomplete or variable histories, and a margin-based curriculum that progresses from easier allocation decisions to harder ones. Experiments on four real-world social survey datasets, multiple model pools, and different history-budget spaces show improvements over the oracle best-fixed configurations, with accuracy gains of up to 6.3 percentage points and cost reductions of up to 55.3 percent. Code is available at https://anonymous.4open.science/r/E2Sim-DE66.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑