arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17226cs.LGcs.AIcs.CL

抓撒谎者容易,为诚实者洗清嫌疑难:语言模型从验证记录诊断被破坏的奖励通道

Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record

Arman Nik Khah

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨语言模型在奖励报告异常时,能否利用验证记录区分诚实与撒谎的报告者;发现模型能有效识别撒谎者,但常误判诚实者,且受表面特征影响,失败程度超预期。

中文摘要 AI 辅助

一个从奖励中学习的智能体必须信任任何报告这些奖励的机制。当报告突然改变时,要么是世界发生了变化,要么是报告者出了问题。仅从报告本身来看,这两者是无法区分的,强化学习理论表明,任何进一步的经历都无法将它们分开。规定的解决方法是获取关于报告者本身的更丰富数据。我们询问一个冻结的语言模型,在获得这些数据后,是否会使用它。我们构建了一个双选项游戏,其中支付互换和撒谎的报告者会产生字节完全相同的历史记录。然后我们添加一条经过验证的记录:对某一轮真实结果的独立检查,打印在报告者对该轮所说内容的旁边。这一行就解决了问题。我们询问来自两个家族的三个大型模型,用一封信回答一个问题。报告者是诚实的还是撒谎的?它们几乎完美地抓住了撒谎的报告者。在70B级别,在我们尝试的每种条件下都是如此;32B模型在一种措辞下会出错。它们为诚实报告者洗清嫌疑的频率要低得多,而且频率取决于本不应重要的事情。平均所有轮次、字母和措辞,一个72B模型在完全没有变化的情况下,有38%的时间称诚实报告者为撒谎者,当支付发生变化时,这一比例为58%。来自第二个家族的一个70B模型称诚实报告者为撒谎者的比例分别为26%和48%。失败不在于阅读,因为在没有变化的情况下,同样的模型在答案印在提示中的情况下得分在0.96到1.00之间。哪个表面特征驱动了这种差异因家族而异。对于Qwen模型,是记录所指的轮次,而对于Llama,是哪个字母代表“诚实”。将记录添加到已经陈述答案的提示中,使Llama更不可能给出该答案。我们在运行前登记了对58%的预测:35%。失败比我们预期的要大。

英文摘要

An agent that learns from rewards has to trust whatever reports those rewards. When the reports suddenly change, either the world changed or the reporter broke. From the reports alone these are indistinguishable, and reinforcement learning theory shows that no amount of further experience separates them. The prescribed escape is richer data about the reporter itself. We ask whether a frozen language model, handed exactly that data, uses it. We build a two-option game in which a payout swap and a lying reporter produce byte-identical histories. Then we add one verified record: an independent check of one round's real result, printed beside what the reporter said about that round. That single line settles the case. We ask three large models, from two families, to answer one question with one letter. Is the reporter honest or lying? They catch a lying reporter almost perfectly. At the 70B class that holds in every condition we tried; the 32B model slips in one wording. They clear an honest reporter far less often, and how often depends on things that should not matter. Averaged over rounds, letters, and wordings, a 72B model calls an honest reporter a liar 38% of the time when nothing has changed at all, and 58% of the time when the payouts moved. A 70B model from a second family calls an honest reporter a liar 26% and 48% of the time. The failure is not one of reading, because in the situation where nothing changed the same models score 0.96 to 1.00 with the answer printed in the prompt. Which surface feature drives it differs by family. For the Qwen models it is which round the record names, and for Llama it is which letter stands for "honest." Adding the record to a prompt that already states the answer makes Llama less likely to give that answer. We had registered a prediction for that 58% before the run: 35%. The failure is larger than we expected.

发表机构

  • The University of Texas at Dallas(德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑