发表机构
SnT, University of Luxembourg(卢森堡大学SnT研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对8个法律RAG系统在两个法律语料库上开展细粒度幻觉分析,发现幻觉普遍存在,不同系统幻觉占比差异大,错误前提问题易引发高幻觉率。
AI 中文摘要
幻觉是检索增强生成(RAG)系统在法律领域面临的重大挑战,无根据的回答可能导致严重后果。为更好地理解该问题,我们对8个法律RAG系统在两个法律语料库(英文的GDPR与法文的某国民法典)中的幻觉行为开展细粒度分析。通过主张级与答案级评估,我们报告了幻觉密度与严重程度,分析了不同问题类别和用户角色下的性能,并在142个由法律专家编写的独立问题集上验证了研究结果。我们的研究表明,幻觉仍普遍存在:表现最佳的系统中,幻觉占比低于10%,而表现最差的系统中这一比例接近一半。我们还发现,包含错误假设、需被拒绝的错误前提问题,在人工编写的问题上会产生高幻觉率。
英文摘要
Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR (in English) and a national civil law (in French). Using claim-level and answer-level evaluation, we report on hallucination density and severity, analyze performance across question categories and user personas, and validate our findings on an independent set of 142 legal-expert-authored questions. Our results show that hallucinations remain pervasive, ranging from less than 10% of responses for the best-performing systems to nearly half in the worst case. We further find that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.