arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2508.13107cs.CLcs.IR

一切为了法律,法律为了一切:面向法律研究的自适应RAG流水线

All for law and law for all: Adaptive RAG Pipeline for Legal Research

  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

Figarri Keisha, Prince Singh, Pallavi, Dion Fernandes, Aravindh Manivannan, Ilham Wicaksono, Faisal Ahmad, Wiem Ben Rim

更新

AI总结:

本文提出一种自适应RAG流水线,通过上下文感知查询转换、开源嵌入及综合评估框架,实现法律研究检索质量媲美专有方法,并提升答案忠实度。

AI中文摘要:

检索增强生成(RAG)通过将大型语言模型(LLM)的输出基于检索到的知识,彻底改变了我们处理文本生成任务的方式。这一能力在法律领域尤为关键。在本工作中,我们引入了一种新颖的端到端RAG流水线,通过三项有针对性的增强改进了以往的基线方法:(i)一种上下文感知的查询转换器,将文档引用从自然语言问题中分离出来,并根据专业性和具体性调整检索深度和响应风格;(ii)使用SBERT和GTE嵌入的开源检索策略,在保持成本效益的同时实现了显著的性能提升;(iii)一个综合性的评估与生成框架,结合RAGAS、BERTScore-F1和ROUGE-Recall来评估不同模型和提示设计下的语义对齐性和忠实度。我们的结果表明,精心设计的开源流水线在检索质量上可以与专有方法相媲美,而自定义的法律接地提示始终比基线提示生成更忠实且上下文相关的答案。综上所述,这些贡献展示了任务感知、组件级调优的潜力,能够为法律研究辅助提供法律接地、可复现且成本效益高的RAG系统。

英文摘要:

Retrieval-Augmented Generation (RAG) has transformed how we approach text generation tasks by grounding Large Language Model (LLM) outputs in retrieved knowledge. This capability is especially critical in the legal domain. In this work, we introduce a novel end-to-end RAG pipeline that improves upon previous baselines using three targeted enhancements: (i) a context-aware query translator that disentangles document references from natural-language questions and adapts retrieval depth and response style based on expertise and specificity, (ii) open-source retrieval strategies using SBERT and GTE embeddings that achieve substantial performance gains while remaining cost-efficient, and (iii) a comprehensive evaluation and generation framework that combines RAGAS, BERTScore-F1, and ROUGE-Recall to assess semantic alignment and faithfulness across models and prompt designs. Our results show that carefully designed open-source pipelines can rival proprietary approaches in retrieval quality, while a custom legal-grounded prompt consistently produces more faithful and contextually relevant answers than baseline prompting. Taken together, these contributions demonstrate the potential of task-aware, component-level tuning to deliver legally grounded, reproducible, and cost-effective RAG systems for legal research assistance.

补充信息

↑