arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ABAI at COLIEE 2026 Task 1: 基于GraphRAG增强元学习的多阶段检索,以及交叉验证到测试差距的事后研究

ABAI at COLIEE 2026 Task 1: Multi-Stage Retrieval with GraphRAG-Enhanced Meta-Learning, and a Post-Hoc Study of the Cross-Validation-to-Test Gap

Minhan Cho, Soyoung Park, Daejin Choi, Jinyoung Han

arXiv 2609.26237首次发表:更新:

发表机构

AlphaBridge; Sungkyunkwan University; National Assembly Research Service; Ewha Womans University(阿尔法桥(AlphaBridge); 成均馆大学; 国会研究服务处; 梨花女子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ABAI系统用于COLIEE 2026判例法检索,采用多阶段流程,并事后分析交叉验证与测试差距,归因于召回上限、分布偏移和阈值校准,提出改进措施。

AI 中文摘要

我们介绍了ABAI系统在COLIEE 2026任务1(判例法检索)中的提交结果,并对其表现不佳的原因进行了受控研究。该任务抑制了被引用的段落本身,这消除了检索器可能依赖的大部分词汇重叠。我们的流程通过四个独立训练的阶段来应对这一问题:基于引用上下文窗口的多视图BM25结合倒数排名融合、神经重排序、来自实体社区和图注意力网络的图特征,以及基于34个特征的LightGBM元学习器。我们最好的运行在官方测试集上达到了F1=0.177,而交叉验证结果为0.311,我们将这一差距归因于召回率上限、时间分布偏移和阈值校准错误。随后我们对这三个因素进行了测试。在无泄漏协议下,阈值迁移损失0.007 F1,决策质量在按时间顺序的四分位数上保持平稳,并且在两个独立的嵌入空间中,官方测试查询与训练流形之间的距离并不比训练查询彼此之间的距离可测量地更远。对遗漏的分解则恰好将遗漏平均分为从未被检索到的候选和检索到但排名低于截断点的候选。针对每一半的补救措施进行测量,BM25长度归一化调整、事件三元组视图和全内容密集融合将top-200召回率提升了3到7个百分点,而引用图特征在去除自引泄漏后,在八个种子上增加了0.014 F1,而按查询截断规则、零样本重排序器替换和日期过滤器则没有帮助。我们还记录了四个评估伪影,每个伪影在协议修正后都逆转了一个结果。

英文摘要

We present the ABAI submission to COLIEE 2026 Task 1, case law retrieval, together with a controlled study of why it underperformed. The task suppresses the cited passages themselves, which removes much of the lexical overlap a retriever would rely on. Our pipeline answers this with four independently trained stages: multi-view BM25 over citation-context windows with reciprocal rank fusion, neural reranking, graph-based features from entity communities and a graph attention network, and a LightGBM meta-learner over 34 features. Our best run reached F1=0.177 on the official test set, against a cross-validated 0.311, and we attributed that gap to a recall ceiling, temporal distribution shift, and threshold miscalibration. We then tested all three. Under leakage-free protocols threshold transfer costs 0.007 F1, decision quality is flat across chronological quartiles, and the official test queries are not measurably farther from the training manifold than training queries are from each other, in two independent embedding spaces. Decomposing the misses instead splits them exactly evenly between candidates never retrieved and candidates retrieved but ranked below the cut. Measuring the remedies for each half, BM25 length-normalisation tuning, an event-triple view, and full-content dense fusion lift top-200 recall by three to seven points, and citation-graph features add 0.014 F1 over eight seeds once own-citation leakage is removed, while per-query cutoff rules, a zero-shot reranker swap, and a date filter do not help. We also document four evaluation artifacts, each of which reversed a result once the protocol was corrected.

Comments13 pages, 3 figures, 9 tables. Extended version of the paper presented at COLIEE 2026 (Workshop on the Thirteenth International Competition on Legal Information Extraction and Entailment), Singapore, June 2026. Code: https://github.com/rabqatab/coliee2026_ABAI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑