发表机构
The University of Texas at Austin; Adobe Research(德克萨斯大学奥斯汀分校; Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较GPT模型与人类审稿人生成的问题对论文改进的效用,发现GPT问题虽多且覆盖广,但引发编辑的比例较低,并揭示长上下文关注可能削弱推理模型的表现。
AI 中文摘要
我们研究了LLMs生成引发编辑的问题的能力,这些问题的答案将改进论文草稿。在来自ICLR和NeurIPS的成对提交版本和最终版本论文的数据集上,我们比较了有无完整论文上下文的GPT模型生成的问题与人类审稿人问题的有用性。GPT生成更多引发编辑的问题,且其问题与更广泛的编辑相关联,并覆盖了比审稿人问题更广泛的编辑内容。然而,GPT问题中引发编辑的比例要小得多。我们的分析证实,自动化问题对作者有益,并突出了一个示例任务,在该任务中,适当关注长上下文会削弱推理模型生成有用输出的能力。
英文摘要
We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.
CommentsAccepted at the DocInsights Workshop @ EMNLP 2026