arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08881cs.AIcs.CL

理论引导的欺骗检测:基于检索增强生成(RAG)的人工智能探索

Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration

David M. Markowitz, Timothy R. Levine

首次发表
浏览论文内容

中文总结 AI 辅助

该研究基于欺骗理论开发7个RAG模型,对比4个大语言模型与基线模型的欺骗判断,发现RAG模型反应偏差更低但准确率无显著差异,理论引导AI判断当前参数下不可靠,优化后或有潜力。

中文摘要 AI 辅助

本研究基于主流欺骗理论开发了7个检索增强生成(Retrieval-Augmented Generation,RAG)模型,并对比了这些模型与基线模型在欺骗判断上的差异。实验采用来自5个已公开欺骗数据集的700条陈述,涉及4个大语言模型(gpt-4o、claude-sonnet-4-6、ollama/llama3、deepseek-v4-flash)和2种运行类型(RAG与基线),共生成39200次欺骗判断。检测准确率与典型人类准确率一致,RAG模型(54.5%)和基线模型(54.6%)的准确率无统计学差异;RAG模型(57.0%)的真相偏差低于基线模型(59.7%),但效应量极小。理论视角对准确率影响不大,却对反应偏差影响显著,反应偏差范围从高度谎言偏差(可验证性方法,32.2%)到高度真相偏差(真相默认理论,88.1%)。内容效应和模型效应进一步调节了结果。当前参数下理论引导的人工智能判断不可靠,但通过补充数据集、模型测试及理论与数据匹配,或可展现潜力。

英文摘要

The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models. Across 700 statements drawn from five published deception datasets, four large language models (gpt-4o, claude-sonnet-4-6, ollama/llama3, deepseek-v4-flash), and two run-types (RAG vs. baseline), a total of 39,200 deception judgments were rendered. Detection accuracies were consistent with typical human accuracies and not statistically different across RAG (54.5%) and baseline models (54.6%). RAG-based models (57.0%) were less truth-biased than baseline models (59.7%), but the effect size was quite small. Theoretical perspective mattered little for accuracy yet mattered substantially for response bias, which ranged from highly lie-biased (the verifiability approach, 32.2%) to highly truth-biased (truth-default theory, 88.1%). Content effects and model effects further moderated the results. Theory-guided AI judgments are unreliable with current parameters, yet they might show promise with additional datasets, model testing, and theory-to-data matching.

↑