发表机构
King’s College London(伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GRAFT方法,通过图蒸馏生成式检索实现科学文献探索,解决文档级检索无法说明关联原因、传统检索局限于已知邻域的问题,在LitWeave语料上取得良好检索效果,且能标注文献关联维度。
AI 中文摘要
科学论文可能因问题、方法、结果或贡献而关联,但文档级检索器会将这些关联压缩为单一相似度分数,无法说明关联原因。仅基于引用和相似度的检索还会将搜索限制在已知内容的邻域,而生成式检索可直接生成文档标识符,实现科学发现所需的探索性检索。我们将论文连接成图,其边由从维度项和引用信号衍生的四个维度(问题、方法、结果、贡献)进行类型标注,并将该图蒸馏为生成式检索器,其标识符为论文自身的维度文本。两种图属性无法通过朴素蒸馏保留:其一,由于每个训练对都是一条边,朴素枚举仅索引84%的语料;覆盖感知蒸馏使每篇论文可通过反向邻域 fallback、最小覆盖阈值和边重要性加权实现可学习性。其二,约束解码保证每个生成的标识符都是有效论文,但无法保证图将其与查询关联;图加权 reciprocal rank fusion 按查询-候选边权重缩放每个候选的排名项,删除无支持的候选。在我们构建的包含11359篇NLP论文的语料LitWeave上,Graft在推理时无需最近邻索引或编码器,即可恢复其图教师的91% Recall@20,且在语料外的查询论文上表现优于图教师;它以0.922的精度复现图自身的维度标签,因此每篇返回论文都会标注其被检索到的维度,而非不透明分数。
英文摘要
Scientific papers may relate by problem, method, result, or contribution, but document-level retrievers collapse these into a single similarity score without saying why they are related. Citation- and similarity-based retrieval alone also confines search to the neighbourhood of what is already known, whereas generative retrieval generates document identifiers directly, enabling the exploratory retrieval that scientific discovery depends on. We connect papers in a graph whose edges are typed by these four facets, derived from facet items and citation signals, and distil it into a generative retriever whose identifiers are the papers' own facet text. Two graph properties do not survive naive distillation. First, because every training pair is an edge, naive enumeration indexes just 84% of the corpus. Coverage-aware distillation makes every paper learnable through a reverse-neighbour fallback, a minimum-coverage threshold, and edge-importance weighting. Second, constrained decoding guarantees that every generated identifier is a valid paper, but not that the graph connects it to the query. Graph-weighted reciprocal rank fusion scales each candidate's rank term by its query-candidate edge weight, dropping unsupported ones. On LitWeave, our constructed corpus of 11,359 NLP papers, Graft recovers 91% of its graph teacher's Recall@20 with no nearest-neighbour index or encoder at inference, and outperforms the graph teacher on query papers outside the corpus. It reproduces the graph's own facet labels at 0.922 precision, so every returned paper arrives labelled with the facet that surfaced it rather than an opaque score.