arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ResearchQA:科学论文中基于引用的问答基准测试

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

Saba Imran, Debanjum Singh Solanky

arXiv 2607.11074首次发表:更新:

发表机构

Khoj Inc.(Khoj公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究引入ResearchQA基准测试,含多领域多类型问答对,用于科学论文基于引用的问答评估。通过特定方法评估八个模型,发现基于引用指标区分度更高,开放权重模型接近最佳封闭模型准确率且延迟更低,还发布了相关资源。

AI 中文摘要

大语言模型越来越多地用于辅助科学阅读,但现有评估方法往往无法检测答案是否有可验证的引用支持。我们引入了ResearchQA,这是一个基准测试,包含来自494篇开放获取论文的6211个单篇论文问答对,涵盖八个领域和四种问题类型:查找、理解、多跳和对抗性。ResearchQA专为基于引用的评估而设计,允许一个主张有多个有效的支持段落,并在源论文不支持答案时奖励有根据的拒绝。我们使用确定性引用匹配器和基于大语言模型的评分评估器,在基于引用的与论文聊天设置中评估了八个领先的封闭和开放权重模型。基于引用的指标比大语言模型评估分数更能清晰地区分系统:各模型的章节覆盖率和引用准确率差异很大,而评估分数则紧密压缩。我们还发现,开放权重模型接近最佳封闭模型的引用准确率,同时每个示例的延迟降低了3到6倍。我们发布了基准测试、评估工具和评估提示。

英文摘要

Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers spanning eight domains and four question types: lookup, comprehension, multi-hop, and adversarial. ResearchQA is designed for citation-grounded evaluation: it permits multiple valid supporting passages for a claim and rewards grounded refusal when the source paper does not support an answer. We evaluate eight leading closed- and open-weight models in a citation-grounded chat-with-paper setting using a deterministic citation matcher and an LLM-based rubric evaluator. Citation-based metrics separate systems more clearly than LLM-evaluator scores: section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed. We further find that open-weight models approach the best closed-model citation accuracy while achieving 3 to 6 times lower per-example latency. We release the benchmark, evaluation harness, and evaluator prompt.

Comments19 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑