arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

欺骗的语义:针对通用领域最先进模型的法律欺骗检测基准测试

Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art

Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb

arXiv 2607.29066首次发表:更新:

发表机构

University of Moratuwa; The Law Office of George C. Lobb(莫拉图瓦大学; 乔治·C·洛布律师事务所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对法律领域的自动欺骗检测,对比了微调Transformer模型与大型语言模型在通用和法律数据集上的表现,发现领域敏感性显著,思维链提示效果逊于直接分类,强调需开发适配法律领域的可解释系统。

AI 中文摘要

欺骗检测对法律诉讼、执法和网络安全具有关键意义。尽管人类判断在准确性和可扩展性上存在局限,但自然语言处理(NLP)提供了数据驱动的替代方案。本文针对法律领域开展基于NLP的自动欺骗检测(ADD)的综述与比较分析,梳理从基于特征的机器学习到大型语言模型(LLM)方法的演进历程。我们在7个数据集(2个法律领域、5个通用领域)上开展统一实证评估,对比6个微调后的Transformer模型与7个LLM,采用4种提示策略。结果显示存在强烈的领域敏感性:微调模型在数据丰富的通用领域表现优异,少样本LLM在低资源法律场景仍具竞争力;思维链提示往往表现不如直接分类。这些发现凸显高风险法律场景中领域适配与可解释系统的必要性。

英文摘要

Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven alternative. We present a survey and comparative analysis of NLP-based Automatic Deception Detection (ADD) focusing on the legal domain, reviewing the evolution from feature-based machine learning to Large Language Model (LLM) approaches. We conduct a unified empirical evaluation across seven datasets (two legal, five general-domain), comparing six fine-tuned transformer models and seven LLMs under four prompting strategies. The results show strong domain sensitivity, with fine-tuned models excelling in data-rich general domains and few-shot LLMs remaining competitive in low-resource legal settings. Chain-of-Thought prompting often underperforms direct classification. These findings highlight the need for domain adaptation and interpretable systems in high-stakes legal contexts.

Comments5 pages paper

Journal refICAIL 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑