arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LitTraceQA:面向科学问答的多阶段溯源与验证基准

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

Xuye Liu, Yimu Wang, Peng Shi, Bo Xue, Xiangrui Ke, Songcheng Cai, Kath Choi, Di Wu, Freda Shi, Krzysztof Czarnecki

arXiv 2608.07370首次发表:更新:

发表机构

University of Waterloo; Amazon; City University of Hong Kong; University of Amsterdam(滑铁卢大学; 亚马逊公司; 香港城市大学; 阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出LitTraceQA基准,用于科学问答的多阶段溯源与验证,提供含多类型证据的问答数据,可评估系统的论文检索、证据溯源及答案准确性,助力生成可验证的科学问答结果。

AI 中文摘要

科学文献正日益成为语言模型、检索增强生成系统及研究助手的知识来源,但回答来自论文的研究问题仅靠流畅生成是不够的。可靠的系统必须识别相关论文、定位支持答案的具体证据,并生成忠实于该证据的响应。我们提出LitTraceQA,一个面向科学论文的文献溯源问答基准。给定一个研究问题和论文元数据集,系统必须返回三个关联输出:规范论文标识符、支持证据位置,以及一种或多种请求格式的答案,包括自由文本、多项选择答案和结构化表格。LitTraceQA针对科学阅读中常见的证据类型:表格、图表、文本片段、方程或算法,以及引用上下文。公开开发集包含55个示例,其中包括26个隐式来源的单篇论文问题和29个多篇论文问题,并提供用于本地验证的黄金论文、证据注释和答案。我们还分析了更大规模的最终注释集合,该集合包含4978条独特问题记录,涉及4859篇独特黄金论文。通过分别评估论文检索、证据溯源和答案准确性,LitTraceQA为科学问答系统提供了一个测试平台,该平台可生成可验证的答案,而非无依据的摘要。

英文摘要

Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent generation. A reliable system must identify the relevant papers, locate the concrete evidence that supports the answer, and produce a response that is faithful to that evidence. We present LitTraceQA, a benchmark for literature-grounded question answering over scientific papers. Given a research question and a metadata pool of papers, a system must return three connected outputs: canonical paper identifiers, supporting evidence locations, and answers in one or more requested formats, including free-form text, multiple-choice answers, and structured tables. LitTraceQA targets evidence types common in scientific reading: tables, figures, text spans, equations or algorithms, and citation contexts. The public development split contains 55 examples, including 26 hidden-source single-paper questions and 29 multi-paper questions, and provides gold papers, evidence annotations, and answers for local validation. We also analyze a larger final annotation collection with 4,978 unique-question records over 4,859 unique gold papers. By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.

CommentsWork in Progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑