arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从宏观社会信号学习模拟个体

Learning to Simulate Individuals from Macro Social Signals

Yining Zhao, Bushi Liu, Haofei Yu, Zhengyang Qi, Shanyong Wang, Chuyue Li, Yuxiang Liu, Jiaxuan You

arXiv 2610.07062首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出macro2mind,利用预测市场价格信号通过GRPO训练语言模型进行显式行为推理,实现零样本用户模拟,并在多个基准上达到最先进性能。

AI 中文摘要

大型语言模型越来越多地被用于模拟个体对新情境的反应,然而这些反应背后的行为推理要么继承自预训练,要么学习自个体层面的标注,这提供了有限的行为多样性,并且对推理本身几乎没有监督。我们提出从预测市场中学习行为推理,预测市场的价格轨迹记录了人群如何大规模地响应现实世界事件。我们引入了macro2mind,它使用GRPO和市场信号训练语言模型。一种社会行为分解使行为推理成为预测的一个显式步骤:模型推断市场参与者的代表性群体,预测每个群体如何解读新闻并更新其信念,推理它们之间的相互作用,并将这些响应聚合成一个价格。一个带有难度感知采样的后见之明遗憾课程将训练集中在后见之明确定的群体显著改善预测的转换上,同时优先考虑当前策略仍可学习的示例。学习到的推理无需进一步训练即可应用于用户模拟。在SWM-Bench上,macro2mind在Polymarket上达到了最先进的定向准确性和相关性。在市场数据上训练后,它零样本迁移到四个用户模拟基准(Humanual、OvertonBench、PRISM和CAD),并在零样本方法中具有竞争力的性能。作为数据生成器使用时,macro2mind还将下游模拟器在未见用户上的准确性提高了15.5个百分点,比其骨干生成的数据高出13.2个百分点。

英文摘要

Large language models are increasingly used to simulate how individuals respond to new situations, yet the behavioral reasoning behind these responses is either inherited from pretraining or learned from individual-level annotations, which offer limited behavioral diversity and little supervision of the reasoning itself. We propose to learn behavioral reasoning from prediction markets, whose price trajectories record how populations respond to real-world events at scale. We introduce macro2mind, which trains a language model with GRPO using market signals. A social behavioral decomposition makes behavioral reasoning an explicit step of forecasting: the model infers representative groups of market participants, predicts how each interprets the news and updates its beliefs, reasons about their interactions, and aggregates these responses into a price. A hindsight-regret curriculum with difficulty-aware sampling focuses training on transitions where hindsight-identified groups substantially improve the forecast while prioritizing examples that remain learnable for the current policy. The learned reasoning applies to user simulation without further training. On SWM-Bench, macro2mind achieves state-of-the-art directional accuracy and correlation on Polymarket. Trained on market data, it transfers zero-shot to four user-simulation benchmarks (Humanual, OvertonBench, PRISM, and CAD) and has competitive performance among zero-shot methods. Used as a data generator, macro2mind also raises a downstream simulator's accuracy on unseen users by 15.5 points, outperforming data generated by its backbone by 13.2 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑