arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型中提示不实回答引发的推理令牌尖峰

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Maverick Morales, Tomáš Dominik, Vermut Gao, Katrina Shirey, Paulius Rimkevičius, Aaron Schurger, Uri Maoz

arXiv 2610.10405首次发表:更新:

发表机构

Chapman University; INSERM U992, Cognitive Neuroimaging Unit, NeuroSpin Center; University of California, Los Angeles; California Institute of Technology(查普曼大学; 法国国家健康与医学研究院U992认知神经影像单元,神经自旋中心; 加州大学洛杉矶分校; 加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨通过推理令牌数量检测大模型不实回答,发现明确提示不实回答会显著增加推理令牌使用,可作为内容无关的检测信号。

AI 中文摘要

监控推理型人工智能模型的思维链仍是检测此类模型中欺骗及其他形式不当行为的关键方法。然而,语义层面的思维链监控依赖于推理轨迹的可读性,以及其是否充分忠实于产生模型行为的底层计算,更不用说其可访问性。此外,越来越多的证据表明,思维链输出可能很快变得难以辨认或不忠实,甚至可能不再可访问。基于认知负荷理论,我们研究了一种较低带宽的信号——生成的推理令牌数量——该信号无需访问推理轨迹的内容。三个具备推理能力的大型语言模型在系统提示指示其如实回答、虚假回答或不考虑真实性回答的情况下,回答了210道多项选择题,涵盖分析性、描述性和规范性推理类型以及道德和非道德领域。在所有三个模型中,指向真实性的回答所产生的推理令牌数量均少于指向谎言和指向不考虑真实性的回答。这些发现表明,明确提示的不实回答策略可以在测试时推理令牌使用上产生稳健的群体层面差异。虽然尚未确立推理令牌计数作为自发欺骗或一般性错位的检测器,但我们的结果证明了一个概念:当原始推理轨迹不可用或不可靠时,它可以作为区分不实与真实模型行为的简单、内容无关的候选信号。未来工作应测试实例层面的检测率、分布外泛化、习得的欺骗策略、隐藏目标以及在对抗压力下的稳健性。

英文摘要

Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.

Comments20 pages, 9 figures, 3 tables. Code: https://github.com/Wakaranaino/token-spike-project ; Data: https://doi.org/10.5281/zenodo.21895296

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑