在向LLM进行民意调查之前:一个审慎诊断框架
Before You Poll with LLMs: A Deliberative Diagnostic Framework
浏览论文内容
中文总结 AI 辅助
提出审慎民调诊断框架,比较人类与LLM在信息干预后的信念转变,发现五个前沿模型均以不同方式失败,表现为反转、过冲或僵化,并揭示自我谄媚特征,建议在信任LLM角色扮演前先进行诊断。
中文摘要 AI 辅助
大型语言模型(LLM)能否像人类一样对新信息进行推理,还是仅仅检索已缓存的观点?这对于硅基抽样至关重要,在该方法中,LLM角色扮演(personas)被用于大规模模拟公众舆论。当前的评估仅测试角色扮演是否持有正确的观点——这是一种静态快照。然而,舆论研究日益依赖于动态保真度:即角色扮演是否像人类在审议过程中那样,根据新论点更新信念。目前尚无基准测试对此进行评估。我们引入了审慎民调诊断框架(Deliberative Polling Diagnostic Framework),该框架比较在相同信息干预后人类与LLM的信念转变。该框架基于审慎民调,能够揭示静态评估中不可见的失败:产生看似合理的党派观点的模型,仍可能错误地呈现这些观点如何变化。我们利用“同一个美国”(America in One Room)的数据(526个角色扮演,72个问题)将该框架应用于五个前沿模型,发现每个模型都以独特的方式失败。GPT-5.1表现出反转:在获得平衡信息后,其角色扮演对对立党派变得更加敌对,而人类则变得不那么敌对。这种反转具有选择性(在外群体问题上为80%,在政策问题上为26%),且在不同党派身份间对称。Gemini 2.0 Flash、Claude Sonnet 4.5和Llama 3.3 70B表现出过冲,即转变方向正确但幅度为人类的5至7倍。DeepSeek V3表现出僵化,变化几乎为零。针对性的消融实验表明,政策内容触发了这些失败,且这些失败具有身份特异性:GPT-5.1在外群体问题上反转,但在内群体问题上过冲;Gemini则呈现相反模式。我们将这一特征称为自我谄媚(self-sycophancy):即模型对其角色扮演的内部刻板印象的遵从,而非基于所提供信息进行推理。我们的框架提供了一个具体协议:在信任LLM角色扮演模拟修正后的信念之前,先运行审慎诊断。
英文摘要
Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opinions can still misrepresent how those opinions change. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), we find that every model fails, each in a unique manner. GPT-5.1 exhibits reversal: its personas become more hostile toward the opposing party after balanced information, while humans become less so. This reversal is selective (80% on outgroup vs. 26% on policy questions) and symmetric across partisan identities. Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting correctly but at 5-7x human magnitude. DeepSeek V3 exhibits rigidity with near-zero change. Targeted ablations reveal that policy content triggers these failures and that they are identity-specific: GPT-5.1 reverses on outgroup questions but overshoots on ingroup; Gemini shows the inverse. We term this signature self-sycophancy: conformity to the model's internal stereotype of the persona rather than reasoning from the information provided. Our framework offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.
发表机构
- Lahore University of Management Sciences(拉合尔管理科学大学)
机构由 AI 辅助整理,请以论文原文为准。