arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于法律信息处理的跨架构大语言模型集成、基于特征的重新排序和检索增强提示

Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing

Amal Saad Alshehri, Nelly Bencomo, Amir Atapour-Abarghouei

arXiv 2607.11400首次发表:更新:

发表机构

Durham University; Jazan University(杜伦大学; 吉赞大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究参与COLIEE 2026五项法律信息处理任务,采用开放权重系统。核心方法包括跨架构模型集成、基于特征重新排序及检索增强提示。在多任务中取得不同成果,如法规蕴含任务准确率高,展示了不同方法在不同任务设置下的有效性。

AI 中文摘要

法律信息处理涵盖检索、蕴含和判决预测等问题,需要在有限监督下进行文本匹配、推理和稳健泛化。我们报告了DU团队参与COLIEE 2026所有五项任务的情况,使用开放权重系统进行法律案例检索、案例蕴含、法规检索与蕴含以及法律判决预测。对于任务3和4,所有模型早于规则要求的2025年7月15日截止日期。在任务4(法规蕴含)中,来自三个家族的九个模型的跨架构集成达到96.3%的准确率,在11个团队的33份提交中排名第一。对于试点任务(侵权预测和理由提取),一个结合五个索赔级模型并使用从索赔预测中得出的特征来完善判决的多视图系统作为非官方提交,TP准确率达到73.1%,RE F1达到68.2%,在TP上高于所有官方参赛作品,在RE上与最高得分匹配。对于任务2(法律案例蕴含),在发布的黄金标签的赛后评估中,仅将提示从单选改为多选,F1从0.343提高到0.555,超过最佳官方提交(F1 = 0.490)。对于任务3(法规检索与蕴含),在赛后分析中,用Qwen3 - 235B和结构化法律推理提示替换蕴含模型,准确率从79.3%提高到91.5%。对于任务1(法律案例检索),一个结合词汇和语义检索以及结构、引用权威和时间特征(共34个)的排序学习系统,F1 = 0.314(在22个团队的54份提交中排名第11)。总体而言,法律信息处理受益于不同任务的不同归纳偏差,跨架构集成、基于特征的重新排序和检索增强提示在不同设置中各显成效。

英文摘要

Legal information processing spans retrieval, entailment and judgment prediction problems, requiring text matching, reasoning and robust generalisation with limited supervision. We report Team DU's participation in all five tasks of COLIEE 2026, using open-weight systems for legal case retrieval, case entailment, statute retrieval and entailment, and legal judgment prediction. For Tasks 3 and 4, all models predate the 15 July 2025 cutoff required by the rules. For Task 4 (statute entailment), a cross-architecture ensemble of nine models from three families achieves 96.3% accuracy, placing first among 33 submissions from 11 teams. For the Pilot Task (tort prediction and rationale extraction), a multi-view system combining five claim-level models and refining the verdict using features derived from the claim predictions achieves 73.1% TP accuracy and 68.2% RE F1 as an unofficial submission, scoring above all official entries on TP and matching the highest on RE. For Task 2 (legal case entailment), changing only the prompt from single- to multi-selection raises F1 from 0.343 to 0.555 in post-competition evaluation on released gold labels, exceeding the best official submission (F1 = 0.490). For Task 3 (statute retrieval and entailment), replacing the entailment model with Qwen3-235B and a structured legal reasoning prompt raises accuracy from 79.3% to 91.5% in post-competition analysis. For Task 1 (legal case retrieval), a learning-to-rank system combining lexical and semantic retrieval with structural, citation authority, and temporal features (34 in total) achieves F1 = 0.314 (rank 11 of 54 submissions from 22 teams). Overall, legal information processing benefits from different inductive biases across tasks, with cross-architecture ensembling, feature-based reranking and retrieval-augmented prompting each proving most effective in different settings.

Comments10 pages. Team DU participation in all five tasks of COLIEE 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑