arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PaperDoctor:面向进行中科学论文的基于证据且可操作的反馈

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou

arXiv 2609.16995首次发表:更新:

发表机构

University of Oxford; National University of Singapore; Stanford University; University of Cambridge; University of Washington(牛津大学; 新加坡国立大学; 斯坦福大学; 剑桥大学; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PaperDoctor是一个智能体框架,通过分层验证和选择性复现实验,为进行中论文提供基于证据且可操作的反馈,将自动评估从评判转向诊断,提升反馈的可审计性。

AI 中文摘要

自主研究智能体正在重塑研究生态系统,但它们也可能让有缺陷的主张大规模进入文献。人类顾问通过仔细、可追踪的反馈在草稿中发现此类问题,然而顾问式评估需要大量人工努力且无法扩展。为了将自动论文评估从评判者转变为诊断者,我们引入了PaperDoctor,一个用于投稿前反馈的智能体框架,具有三项关键创新。首先,一个整体分层框架通过三层评估写作、排版、参考文献、代码、理论、先前工作和实验:L1表面筛选,L2类型化验证器将每个主张路由到适当的证据,以及L3复现器按优先级重新运行实验。其次,每个发现包含一个观察、一个指向具体证据(如句子、方程或代码行)的指针,以及一个修改建议,使批评可审计且可操作。第三,PaperDoctor根据主张重要性和计算预算选择性地重建并重新运行实验,揭示仅从手稿中不可见的可复现性差距和定量局限性。我们在30篇进行中论文上评估PaperDoctor,获得70.6%的一致性和所有正面的整体评分,并在涵盖机器学习、自然科学和社会科学的40篇手稿上评估,包括带代码的人类和AI撰写论文。总体而言,PaperDoctor产生比人类和其他智能体审稿人更可审计的反馈,通过设计将批评与具体建议配对,并补充了人类审稿人经常忽视的维度。我们还开发了一个交互式界面,让作者浏览基于其论文的发现。PaperDoctor将自动论文评估重新定义为诊断而非裁决,朝着更严谨的AI辅助科学发现的AI顾问迈出了具体一步。

英文摘要

Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessment from a judge to a diagnostician, we introduce PaperDoctor, an agent framework for pre-submission feedback with three key innovations. First, a holistic hierarchical framework evaluates writing, layout, references, code, theory, prior work, and experiments through three layers: L1 surface screening, L2 typed verifiers that route each claim to the appropriate evidence, and L3 reproducers that rerun experiments by priority. Second, each finding contains an observation, a pointer to specific evidence such as a sentence, equation, or code line, and a revision suggestion, making critiques auditable and actionable. Third, PaperDoctor selectively rebuilds and reruns experiments based on claim importance and compute budget, surfacing reproducibility gaps and quantitative limitations that are invisible from the manuscript alone. We evaluate PaperDoctor on 30 in-progress papers, yielding 70.6% agreement and all positive holistic scores, and on 40 manuscripts across machine learning, natural science, and social science, covering human- and AI-authored papers with code. Overall, PaperDoctor produces more auditable feedback than human and other agentic reviewers, pairs critiques with concrete suggestions by design, and complements dimensions often overlooked by human reviewers. We also develop an interactive interface that lets authors browse findings grounded in their paper. PaperDoctor reframes automated paper assessment as diagnosis rather than verdict, taking a concrete step toward AI advisors for more rigorous AI-assisted scientific discovery.

CommentsWebsite: http://paperdoctor.github.io/ Github: https://github.com/QinghongLin/paperdoctor

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑