发表机构
Massachusetts Institute of Technology; Flybits Labs; Toronto Metropolitan University(麻省理工学院; Flybits实验室; 多伦多都会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出贝叶斯编年史智能体(BCA),通过分离信念与表达、用贝叶斯步骤更新立场并引入先验强度参数,实现可控意见动态,匹配FJ模型并提升可审计性。
AI 中文摘要
在社会模拟中,LLM智能体通过上下文隐式地修正其意见:智能体对说服的开放程度既无法指定也无法验证,集体结果继承了模型的训练先验。我们引入了贝叶斯编年史智能体(BCA),这是一个最小信念层,将智能体“相信什么”与“如何表达”分离。每个立场是一个概率,每听到一次话语便通过一次贝叶斯步骤更新。一个先验强度参数$\kappa$编码了固执程度,其建模参考了Friedkin-Johnsen(FJ)意见动态中的角色。我们随后扫描该参数,按需生成意见动态的三种典型机制(共识、持续分歧、承诺少数派影响),其中持续分歧与FJ闭式不动点在$R^2=0.93$–$0.99$下匹配。我们进一步表明,规定的$\kappa$在语言往返后仍可恢复,在所有四个模型中实现完美的秩次恢复。显式信念也使模拟可审计:该层揭示了端到端模拟会静默吸收的系统性逐模型立场偏差。
英文摘要
LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior. We introduce Bayesian Chronicle Agents (BCA), a minimal belief layer separating what an agent believes from how it speaks. Each stance is a probability, updated by one Bayesian step per utterance heard. A single prior-strength parameter $κ$ encodes stubbornness, modeled after its role in Friedkin--Johnsen (FJ) opinion dynamics. We then sweep this parameter to yield three canonical regimes of opinion dynamics on demand (consensus, persistent disagreement, committed-minority influence), with persistent disagreement matching the FJ closed-form fixed points at $R^2\!=\!0.93$--$0.99$. We further show that prescribed $κ$ remains recoverable after the language round-trip, with perfect rank-order recovery across all four models. Explicit belief also makes simulation auditable: the layer surfaces systematic per-model stance biases that end-to-end simulation would silently absorb.
CommentsAccepted to The 2nd Workshop for Research on Agent Language Models (REALM) at EMNLP 2026