发表机构
Hong Kong University of Science and Technology; University of Southern California; University of Sydney; University of Hong Kong(香港科技大学; 南加州大学; 悉尼大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出生成-提取流程模型,通过立场保留率(SPR)量化AI中介沟通中的立场漂移,发现九个LLM在112个命题上SPR均低于0.7,并识别极化、偏离中立和翻转三种漂移模式,其中为GPT-5.4添加中等推理努力可将SPR提升至0.775,揭示了AI沟通的保真度差距。
AI 中文摘要
大型语言模型(LLMs)越来越多地介入人类沟通,从起草电子邮件到总结科学报告,但它们是否忠实保留说话者的立场在很大程度上仍未得到检验。我们将AI中介沟通建模为一个两步生成-提取流程:一个LLM从指定立场生成论点,第二个LLM从该论点中提取立场。我们将该流程表示为在五个李克特式立场类别上的概率状态转移,并将立场保留率(SPR)定义为提取立场与初始立场匹配的平均概率。在112个辩论命题中,默认配置下测试的九个LLM均未超过0.7的SPR。三种漂移模式占据了大部分漂移:极化、偏离中立和翻转。在测试的缓解策略中,包括上下文学习、选项打乱的多重提取、断言和反思,只有为GPT-5.4的反思提示添加中等推理努力显著提高了SPR至0.775,但极化仍是最大模式,占转移质量的0.119。对单一命题的探索性人类提取比较表明,漂移在生成和提取阶段都会出现。这些结果指出了AI中介沟通中的保真度差距,对新闻业、政策审议、科学传播以及其他观点性信息通过语言模型传递的领域具有影响。
英文摘要
Large language models (LLMs) increasingly mediate human communication, from drafting emails to summarizing scientific reports, yet whether they faithfully preserve a speaker's position remains largely untested. We model AI-mediated communication as a two-step generation-extraction pipeline: one LLM produces an argument from a specified stance, and a second LLM extracts the stance from that argument. We represent the pipeline as a probabilistic state transition over five Likert-type stance categories and define the stance preservation rate (SPR) as the average probability that the extracted stance matches the initial stance. Across 112 debate propositions, none of the nine LLMs tested exceeded an SPR of 0.7 under the default configuration. Three drift patterns accounted for most of the drift: polarization, deviation from neutrality, and flipping. Among the mitigation strategies tested, including in-context learning, multiple extraction with shuffled options, assertion, and reflection, only adding medium reasoning effort to a reflection prompt for GPT-5.4 substantially improved the SPR, to 0.775, yet polarization remained the largest pattern, with 0.119 of the transition mass. An exploratory comparison with human extraction on a single proposition suggests that drift arises at both the generation and the extraction stage. These results point to a fidelity gap in AI-mediated communication, with implications for journalism, policy deliberation, scientific communication, and other domains where opinion-laden messages pass through language models.