ReGround:将审稿人评论锚定到多模态证据
ReGround: Grounding Reviewer Comments in Multimodal Evidence
- TU Darmstadt(达姆施塔特工业大学)
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态长文档中审稿人评论难以锚定到证据的问题,提出ReGround数据集,将锚定视为检索任务,发现全文检索效果差、证据类型推断是瓶颈,多模态证据提供互补信号。
AI中文摘要:
审稿人评论自然与所审论文的特定部分相关,然而由于多模态文档篇幅较长,将这些评论锚定到其底层证据十分困难。现有基准未涵盖这一场景,且主要聚焦于明确的、信息检索型查询。我们提出了ReGround,一个用于审稿人评论锚定的大规模数据集,该数据集将10,267条审稿人评论与来自3,656篇论文原始匿名投稿中的16,274条证据相关联。我们基于一个简单的观察:作者的反驳意见中常包含对投稿中用于回应审稿人评论的内容的明确引用,这提供了高精度的标注来源。我们将锚定任务视为检索任务,并评估了多种检索方法。结果表明,对整篇论文内容进行检索的效果不佳,证据类型推断是主要瓶颈,且多模态证据提供了文本单独无法提供的互补信号。我们的数据集揭示了审稿人评论锚定是科学文档理解中一个困难且具有实际重要性的问题。
英文摘要:
Reviewer comments naturally relate to specific parts of the reviewed paper, yet grounding these comments to the underlying evidence is difficult due to long multimodal documents. Existing benchmarks do not capture this setting and largely focus on explicit, information-seeking queries. We introduce ReGround, a large-scale dataset for reviewer comment grounding that links 10,267 reviewer comments to 16,274 evidence in the original anonymous submission of 3,656 papers. We build on a simple observation: author rebuttals often include explicit references to content of the submission used to address reviewer comments, providing a high-precision annotation source. We cast grounding as a retrieval task and evaluate a wide range of retrieval methods. Results show that retrieval over the entire paper content performs poorly, evidence-type inference is a major bottleneck, and multimodal evidence provides complementary signals that text alone misses. Our dataset exposes grounding reviewer comments as a difficult and practically important problem for scientific document understanding.