arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

L-MAD:法律推理中多智能体辩论结构的系统评估

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

Tan-Minh Nguyen, Hoang-Trung Nguyen, Huu-Dong Nguyen, Dinh-Truong Do, Thi-Hai-Yen Vuong, Le-Minh Nguyen

arXiv 2607.09099首次发表:更新:

发表机构

Japan Advanced Institute of Science; VNU University of Engineering(日本先进科学研究院; 越南工程大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在法律领域探索多智能体辩论(MAD)框架有效性,引入L-MAD框架评估不同结构和方法,通过分配专家角色提升效果,分析辩论规模发现权衡,明确了在法律推理中部署协作多智能体系统的边界与安全边际。

AI 中文摘要

虽然多智能体辩论(MAD)框架在一般推理中显示出巨大潜力,但在高度结构化、知识密集的法律领域其有效性仍未得到充分探索。在这项工作中,我们引入法律多智能体辩论(L-MAD)框架,以系统评估法律文本蕴含中的不同辩论结构和聚合方法。通过为多个智能体分配不同的专家角色,L-MAD比强大的单智能体基线提高了8%。分析辩论规模显示出明显的权衡:增加智能体数量可减少不一致性并提高准确性,而延长讨论轮次会引发有害的“过度审议漂移”,即智能体强化彼此的错误。最终,我们的发现勾勒出在高风险法律推理环境中部署协作多智能体系统的实际边界和安全边际。

英文摘要

While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains under-explored. In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within Legal Textual Entailment. By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8\%. Furthermore, analyzing how debate scales reveals a clear trade-off: increasing the agent population reduces inconsistency and improves accuracy, whereas extending discussion rounds induces a detrimental \textit{over-deliberation drift} where agents reinforce each other's mistakes. Ultimately, our findings outline the practical boundaries and safety margins of deploying collaborative multi-agent systems in high-stakes legal reasoning environments.

CommentsOutstanding paper in the AI4Law Workshop at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑