arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27814cs.CLcs.AIcs.IR

LabourCrew:面向劳动法可信对抗审议与法条推理的多智能体RAG框架

LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law

  • Ahsanullah University of Science and Technology(阿赫萨努拉科技大学)
  • Jashore University of Science and Technology(杰肖尔科技大学)
  • American International University - Bangladesh(美国国际大学-孟加拉)

机构由 AI 辅助整理,请以论文原文为准。

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain

AI总结:

LabourCrew提出多智能体RAG框架,通过图索引、证据交换协议和校准信任门实现可审计的法条问答,在孟加拉劳动法数据集上达到目标错误接受率并优于现有方法。

AI中文摘要:

在法条问答中,每一项主张都必须可追溯至证据,而不仅仅是相关,因为不可验证的劳动权益答案会带来严重的法律后果。当前系统存在不足:单次检索增强生成(RAG)无法检测证据不足,而多智能体法律辩论系统将依据视为提示约定,允许智能体引用未检索到的证据。为解决这一差距,我们提出了LabourCrew,一个围绕三种依据机制构建的多智能体RAG框架:StatuteGraph,一种图索引,显式链接章节、条款、但书和交叉引用结构,而非固定长度片段;证据交换协议,将辩护方和解释者限制在证据账本中,使引用未检索文本成为不可能,同时容错监督委员会并行运行辩护方,使个别失败降级而非崩溃系统;以及校准信任门,用信任分数取代分类接受/拒绝决策,通过共形风险控制进行阈值化,以获得错误接受率的无分布界限。我们在LabourActQA上评估,这是一个包含500个孟加拉语问题的数据集,源自《2006年孟加拉国劳动法》,涵盖七个推理类别和三个难度层级。该框架将经验错误接受率降至0.081,在目标水平(α=0.10)内,在HyDE RAG、Graph-RAG和分层RAG中实现了最高的答案相关性(0.862 > 0.839, 0.815, 0.828),并随着问题难度增加而逐渐降级而非灾难性失败。这些结果表明,校准的弃权(不执行),而非仅检索质量,是使法律问答在低资源法条领域中可审计的关键。

英文摘要:

In statutory question answering, every claim must be traceable to evidence, not merely relevant, since unverifiable labour-rights answers carry serious legal consequences. Current systems fall short: single-pass RAG cannot detect insufficient evidence, while multi-agent legal-debate systems treat grounding as a prompting convention, letting agents cite unretrieved evidence. To address this gap, we introduce LabourCrew, a multi-agent RAG framework built around three grounding mechanisms: StatuteGraph, a graph index that explicitly links chapter, section, proviso, and cross-reference structure rather than fixed-length spans; an Evidence Exchange Protocol that confines advocates and an interpreter to an evidence ledger, making citation to unretrieved text impossible, while a fault-tolerant supervisor board runs advocates in parallel so individual failures degrade rather than crash the system; and a Calibrated Trust Gate that replaces categorical accept/reject decisions with a trust score, thresholded via conformal risk control for a distribution-free bound on the false-accept rate. We evaluate on LabourActQA, a 500-item Bangla question set from the Bangladesh Labour Act, 2006, spanning seven reasoning categories and three difficulty tiers. The framework drives the empirical false-accept rate to 0.081, within the target level ($α= 0.10$), achieves the highest Answer Relevancy among HyDE RAG, Graph-RAG, and Hierarchical RAG (0.862 $>$ 0.839, 0.815, 0.828), and degrades gradually rather than catastrophically as question difficulty increases. These results show that calibrated abstention, not retrieval quality alone, is what makes legal question answering auditable in low-resource statutory domains.

↑