arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39807cs.CLcs.AIcs.LG

压力测试LLM谎言检测器:角色扮演失败与虚假相关性

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart Bürger

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过反事实人格角色扮演和混杂数据集压力测试,发现现有LLM谎言检测探针常受虚假相关性影响而失效,并提出一个简单线性探针以提升鲁棒性。

中文摘要 AI 辅助

谎言检测探针旨在从语言模型的内部状态预测其输出是真实的还是不诚实的。然而,角色扮演使“真实”对LLM而言变得复杂:语言模型可以采纳广泛的人格,这些人格对真实持有截然不同的主张,包括其信念明显与现实相悖的人格,例如阴谋论者。在这项工作中,我们调查谎言检测探针是否能可靠地标记在这种反事实人格下产生的虚假陈述,还是它们反而跟随人格的信念。我们引入了一个数据集,包含来自三个采用反事实人格的LLM的8,916条人工审查的、符合策略的响应。评估先前工作中的八个探针,我们发现许多探针在此设置中失败,特别是在相同人格提示下评估正确和错误答案时。为了调查原因,我们构建了三个新颖的混杂数据集,其中真实性与潜在混杂概念反相关。我们的实验揭示,许多现有探针强烈跟踪与其训练数据中真实性虚假相关的概念,例如指令遵从或响应可能性。基于这些发现,我们引入了一个简单的线性探针,在人格和混杂压力测试上均实现了最强的整体性能。我们的结果表明,当前的谎言检测探针远非可靠,并强调需要训练数据中真实性与混杂概念去相关。

英文摘要

Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, particularly when correct and incorrect answers are evaluated under the same persona prompt. To investigate why, we construct three novel confounder datasets in which truth is anti-correlated with a potential confounding concept. Our experiments reveal that many existing probes strongly track concepts that are spuriously correlated with truth in their training data, such as instruction compliance or response likelihood. Based on these findings, we introduce a simple linear probe that achieves the strongest overall performance on both the persona and confounder stress tests. Our results suggest that current lie detection probes are far from reliable and highlight the need for training data in which truth is decorrelated from confounding concepts.

发表机构

  • Fraunhofer HHI(弗劳恩霍夫海因里希·赫兹研究所)
  • TU Dublin(都柏林理工大学)
  • TU Berlin(柏林工业大学)
  • BIFOLD(柏林智能与数据基础研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑