arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20479cs.AIcs.CL

超越说谎者基准:谎言类型、深度和稀疏性对大语言模型中欺骗检测的影响

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

  • University of Bonn(波恩大学)
  • Lamarr Institute for ML and AI(拉马尔机器学习与人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

Amr Moustafa, Max Feser, Florian Mai

AI总结:

研究大语言模型欺骗检测问题,通过扩充数据对谎言类型等因素影响检测性能进行系统研究,发现最佳表示深度依赖数据集,更具表达力探测器提升有限,稀疏特征与密集状态表现类似,强调训练数据和谎言类型选择对检测性有显著影响。

AI中文摘要:

训练探测器以检测大语言模型中的欺骗性输出仍是一个未解决的问题。近期工作表明检测探测器在域外场景中表现不佳,在一种谎言类型上的训练不能很好地转移到涉及其他类型谎言的欺骗场景。本文对各种因素如何影响检测性能进行了系统研究,包括表示深度、探测器表达能力、稀疏特征表示和训练数据的谎言类型。为此,用包含多种欺骗类型的补充数据集扩充标准基准训练数据。分析七种探测器类型的这些因素,实验结果表明最佳表示深度高度依赖数据集,更具表达能力的探测器仅比线性基线有选择性地提升,稀疏自动编码器特征与密集隐藏状态表现相似。最终证明训练数据和谎言类型的选择会显著改变可检测性,突出欺骗检测是一个高度依赖表示的问题。

英文摘要:

Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not transfer well to deception scenarios involving other types of lies. In this work, we conduct a systematic study on how various factors impact detection performance: representation depth, probe expressivity, sparse feature representations, and the lie typology of the training data. To this end, we augment standard benchmark training data with a supplementary dataset containing diverse types of deception, including fabrication, omission, and exaggeration examples. Analyzing these factors across seven probe types, our experimental results show that the optimal representation depth is highly dataset-dependent, more expressive probes provide only selective gains over linear baselines, and sparse autoencoder features perform similarly to dense hidden states. Ultimately, we demonstrate that the choice of training data and lie typology substantially changes detectability, highlighting that deception detection is a highly representation-dependent problem.

补充信息

↑