arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从欺骗性输出到欺骗性机制:语言模型欺骗研究的因果框架

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Yakov Pyotr Shkolnikov

arXiv 2609.04166首次发表:更新:

AI 中文总结

该研究针对语言模型欺骗中混淆行为与机制的问题,提出因果分类框架,经实验发现看似欺骗的行为可能无对应机制,接收者信息状态会因果影响欺骗偏好,且行为证据不能确立模型代理地位。

AI 中文摘要

关于语言模型欺骗的研究和新闻报道越来越多地将类似人类的心理状态概念归因于语言模型,这类说法可能会模糊看似欺骗的行为与实际欺骗机制之间的界限。我们引入一种因果分类法,区分先验承诺与回溯报告、模型偏好与已实现输出、虚假偏好与对误导接收者效用的敏感性,以及欺骗性行为与产生该行为的目标或策略的来源。我们在两个开放权重模型系列中测试这些区分,在受控猜谜游戏和股票交易实验中发现,看似欺骗的行为可能在没有对应所提出机制的情况下产生,而其他干预措施则提供了直接证据,表明接收者的信息状态可因果性地影响欺骗偏好。这些结果表明,欺骗性行为可作为欺骗机制的证据,但即便有此类机制的证据,也不能确立模型在欺骗中的代理地位。

英文摘要

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑