发表机构
Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LexAgentHallu是一个法律智能体幻觉分层基准,通过3414个实例和双层分类法,诊断多步轨迹中幻觉的程度与方式,揭示正确答案-错误推理效应。
AI 中文摘要
随着大型语言模型越来越多地被部署为工具增强型法律智能体,它们引入了智能体幻觉,其中工具调用和推理错误会级联成捏造的裁决和错误引用的权威。然而,现有的法律基准仅评估具有结果级指标的单轮问答,而智能体幻觉基准缺乏法律特定的诊断能力。两者都无法回答法律智能体在其轨迹中幻觉的程度和方式。为解决这些局限性,我们引入了LexAgentHallu,一个法律智能体幻觉基准,旨在评估法律智能体在多步轨迹中失败的程度和方式。通过一个四阶段专家参与流程构建,LexAgentHallu包含17个法律类别和6种任务类型的3414个实例。每个实例在7个高级类别和27个细粒度子类的双层幻觉分类法下进行标注,涵盖实质性错误和智能体程序性失败。我们进一步设计了细粒度指标,量化每个失败沿智能体执行路径发生的程度并定位其方式。我们对18个专有和开源智能体的评估揭示了“正确答案-错误推理”效应,并发现幻觉子类聚集而非分散,形成不同的智能体框架、法律任务和类别特征。这些发现对结果级评估不可见,验证了LexAgentHallu在评估法律智能体幻觉方面的诊断能力。
英文摘要
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414 instances across 17 legal categories and 6 task types. Each instance is annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses, covering both substantive errors and agent-procedural failures. We further design fine-grained metrics that quantify to what extent and localize how each failure occurs along an agent's execution path. Our evaluation across 18 proprietary and open-source agents uncovers a Right-Answer-Wrong-Reason effect and reveals that hallucination subclasses cluster rather than scatter, forming distinct agentic framework, legal task, and category profiles. These findings, invisible to outcome-level evaluation, validate the diagnostic power of LexAgentHallu for evaluating agentic hallucination in law.
CommentsEMNLP 2026 Main
Journal refEMNLP 2026 Main