发表机构
Tsinghua University; Quancheng Laboratory; Renmin University of China(清华大学; 清华实验室; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对法规检索中口语化查询与正式法规语言的差距问题,提出GCSR生成式法规检索框架,通过多粒度结构化文档ID和多任务训练策略,将其转化为序列生成问题,实验证明该框架优于基线,凸显其在法律信息检索及推理任务中的有效性与潜力。
AI 中文摘要
法规检索是法律信息检索中的一项基本任务,但现有方法难以弥合口语化法律查询与正式法规语言之间的差距。本文提出了GCSR,一种生成式法规检索框架,将法规检索重新表述为序列生成问题,并将法规知识内化到生成模型中。具体而言,提出了一种编码法律层次和语义信息的多粒度结构化文档ID,以及多任务训练策略。实验表明,GCSR始终优于强大的稀疏、密集和法律领域基线。结果证明了生成式检索对法规检索的有效性,并突出了其在更广泛的法律信息获取和下游法律推理任务中的潜力。
英文摘要
Statute retrieval is a fundamental task in legal information retrieval, yet existing approaches struggle to bridge the gap between colloquial legal queries and formal statutory language. In this paper, we propose GCSR, a generative statute retrieval framework that reformulates statute retrieval as a sequence generation problem and internalizes statutory knowledge into a generative model. Specifically, we propose a multi-granularity structured docid that encodes legal hierarchy and semantic information, together with a multi-task training strategy. Experiments show that GCSR consistently outperforms strong sparse, dense, and legal-domain baselines. Our results demonstrate the effectiveness of generative retrieval for statute retrieval and highlight its potential for broader legal information access and downstream legal reasoning tasks.