AI 中文总结
针对文献 grounded 问答,提出 PageRecall 系统,发现证据定位受检索限制,通过展示整篇论文替代页面选择,提升 gold 页面召回率至 100%,并解析参考文献支持位置型问题。
AI 中文摘要
我们描述了用于 LitTraceQA(GroundLM @ EMNLP 2026)的系统:给定一个研究问题,从 27,487 篇论文池中检索相关论文,引用答案所在的页面以及表格或图,并按要求的格式作答。我们的主要发现是,证据 grounding 受限于检索,而非阅读。页面选择器将标注者的页面(我们称之为 gold 页面)放在定位证据的模型面前,仅约一半时间(52.6% 的 gold 页面召回率)成功,而该模型在给定页面时,在其发出的 48 个定位器中,有 45 个引用了正确的页面(94%)。当页面缺失时,它很少说明:在 45 个此类案例中,14 次未返回任何内容,24 次返回错误页面,7 次返回正确页面,因此流水线静默失败的频率几乎是显式失败的两倍。由于失败在于从未显示正确页面,解决方案是停止选择:每篇检索到的论文都适合模型的上下文,因此我们展示整篇论文。页面排序仅作为过长论文的备用方案存在,而测试分割中的论文均未超过长度限制,gold 页面召回率在我们能解析的论文上达到 100%。另外,通过位置而非内容识别目标的问题(如“第 24 篇参考文献的第一作者”)由解析而非检索处理:我们将参考文献解析为可寻址列表,这也提供了证据度量所评分的标识符。最终系统在保留测试分割上得分为 0.762 的论文 F1、0.441 的证据 F1 和 0.920 的多选准确率。由于流水线依赖无种子控制的封闭模型,我们发布了一个验证工具,用于对照已提交的工件验证论文的核心主张。
英文摘要
We describe our system for LitTraceQA (GroundLM @ EMNLP 2026): given a research question, retrieve the relevant papers from a pool of 27,487, cite the page and the table or figure where the answer lives, and answer in a requested format. Our main finding is that evidence grounding is limited by retrieval, not by reading. The page selector put the annotator's page, which we call the gold page, in front of the model that locates evidence only about half the time (52.6% gold-page recall), while that model, given the page, cited the right one in 45 of the 48 locators it emitted (94%). When the page was missing it rarely said so: of 45 such cases it returned nothing 14 times, a wrong page 24 times, and a correct page 7 times, so the pipeline failed quietly almost twice as often as it failed visibly. Since the failure was that the right page was never shown, the fix is to stop choosing: each retrieved paper fits in the model's context, so we show it whole. Page ranking survives only as a fallback inside papers too long to fit, which no test-split paper was, and gold-page recall reaches 100% on the papers we can parse. Separately, questions that identify their target by position rather than content, such as "the first author of the 24th reference", are served by parsing rather than retrieval: we resolve the bibliography into an addressable list, which also supplies identifiers the evidence metric scores. The final system scores 0.762 paper $F_1$, 0.441 evidence $F_1$ and 0.920 multiple-choice accuracy on the held-out test split. Because the pipeline depends on a closed model without seed control, we release a harness that verifies the paper's central claims against committed artifacts.
CommentsAccepted at the 1st Workshop on Grounding Language Models (GroundLM 2026), co-located with EMNLP 2026. 9 pages. System description for the LitTraceQA shared task (team Everest)