无触发器的错误信息:从事实性回答到下游决策
Misinformation Without Triggers: From Factual Answers to Downstream Decisions
浏览论文内容
中文总结 AI 辅助
研究发现语言模型在无触发器数据投毒下,事实性回答正确但下游决策仍受虚假内容影响,存在审计差距,正确事实不保证正确决策。
中文摘要 AI 辅助
语言模型从网络文档中学习,其中一些是虚假的,虚假内容可以到达模型对事实性问题的回答以及使用该回答的摘要和决策。大多数数据投毒研究在训练数据中添加触发器并在提示中激活它。虚假文档也可以在没有触发器的情况下改变事实性回答,但我们不知道直接回答是否能预测决策。在这项工作中,我们追踪虚假内容超越回答本身,发现直接探针报告的内容与模型随后采取的行动之间存在一个“审计差距”,在受控决策任务“猜首都”中,将虚假训练与匹配的真实对照进行比较,其中固定解码器将事实性回答转换为计分卡片选择,并针对来自Facebook帖子中关于2019-20年澳大利亚丛林大火的误导性说法。在剂量为1,000的八个模型中,直接注入选择率达到95.8-100%,而注入的游戏选择相对于匹配的真实训练增加了1.7-14.4个百分点。这种差距也以另一种方式存在。通过直接探针的事实仍然推动决策走向注入的回答,游戏准确性下降幅度超过这些选择所能解释的。真实性的纠正使事实回归,但基于事实构建的决策并未回归。我们随后研究现实世界的丛林大火案例,在虚假帖子训练的模型即使失去夸大的计数,仍表示人们因纵火被捕,并且在计数-措辞因子设计中,即使训练计数保持在24,误导性的纵火措辞也会产生纵火断言。简而言之,一个正确的事实性回答并不能保证正确的决策,失去注入的数字并不能消除误导性的故事。
英文摘要
Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a misleading claim from Facebook posts about the 2019--20 Australian bushfires. Across eight models at dose 1,000, direct injected-choice rates reach 95.8--100\%, while injected game choices increase by 1.7--14.4 percentage points over matched truthful training. The gap runs the other way too. Facts that pass the direct probe still push decisions toward the injected answer, and game accuracy drops further than those choices explain. Truthful correction brings the fact back but not the decisions built on it. We then look into the real-world bushfire case, models trained on the false posts say that people were arrested for arson even when they lose the inflated count, and in a count-by-wording factorial the misleading arrest wording produces arrest assertions even when the training count stays at 24. In simpler terms, \textbf{a correct factual answer does not guarantee a correct decision, and losing the injected number does not remove the misleading story}.
发表机构
- University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。