发表机构
Tilburg University; University College London(蒂尔堡大学; 伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过元分析发现自动言语欺骗检测存在70%-75%的持续准确率上限,采用嵌入技术和大语言模型未提升性能,方法学质量对准确率的影响大于模型复杂度。
AI 中文摘要
为克服人类言语欺骗检测的局限性,人们提出了自动方法,但各学科间的相关证据仍零散。本研究系统综述了25年的研究(289份报告、6136个分类模型),并对97个数据集中嵌套的3653个模型进行元分析。合并准确率为74.4%(95%置信区间:71.2%-77.4%),存在显著异质性。准确率更多由方法学质量(真实标签、数据源、类别平衡、评估程序)而非模型复杂度驱动:采用嵌入技术和大语言模型并未转化为预测性能提升。仅12.46%的报告使用了具有可验证真实标签的数据,仅23.96%的模型在独立数据上评估。合并准确率与手动方法的元分析结果一致,表明存在70%-75%的准确率上限,当前研究惯例不太可能突破该上限。
英文摘要
Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines. We systematically reviewed 25 years of research (289 reports, 6,136 classification models) and meta-analyzed 3,653 models nested within 97 datasets. Pooled accuracy was 74.4% (95% CI: 71.2%-77.4%) with substantial heterogeneity. Accuracy was driven by methodological quality (ground truth, data source, class balance, evaluation procedure) more than by model complexity: the adoption of embeddings and large language models has not translated into improved predictive performance. Only 12.46% of reports used data with verifiable ground-truth, and only 23.96% of models were evaluated on independent data. The pooled accuracy aligns with meta-analyses of manual approaches, suggesting a ceiling of 70-75%, unlikely to be lifted by current research conventions.