arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27228cs.CLcs.AIcs.DL

AI辅助的开源软件提交预评审:来自BOSC 2026的经验报告

AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026

Tazro Ohta, Nomi L. Harris, Seth Carbon

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对BOSC 2026投稿激增问题,开发bosc-pre-review与Runabilly辅助评审,调查显示该预评审工具对志愿评审人员有用,需人工复核结论。

中文摘要 AI 辅助

大多数会议依赖投稿的同行评审,但随着生成式AI让准备投稿材料变得前所未有的容易,部分会议正面临投稿量激增的问题。我们希望探究生成式AI能否通过预评审摘要的特定标准,为会议的志愿评审人员提供帮助。生物信息学开源会议(BOSC)具备开展相关实验的有利条件,因为评审人员已采用详细的评审标准,从多个维度评估投稿摘要,包括项目相关代码或其他内容的公开可用性(开放性)、有效的开源许可证,以及“可运行性”(项目的下载、构建和运行难度,是衡量可复用性的重要指标)。针对2026年BOSC,我们开发了bosc-pre-review,这是一个评估六项评审标准的智能体技能,还开发了Runabilly,它会在一次性Docker容器中构建并测试每个项目以保障安全。AI仅收集证据提交给评审人员,所有关于摘要录用的决策均由人工做出。评审期结束后,我们对评审人员进行调查以确定他们对预评审的评价,多数受访者表示预评审有用,但他们更倾向于核对AI的结论而非不加质疑地接受AI结果。

英文摘要

Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences are seeing an overwhelming surge of submissions. We wanted to see if generative AI could help our conference's volunteer reviewers by pre-reviewing abstracts for certain criteria. The Bioinformatics Open Source Conference (BOSC) was well-positioned to experiment with this, as we already had a detailed rubric used by reviewers to evaluate submitted abstracts on multiple criteria, including openness (public availability of the code or other content associated with the project), valid open source license, and "runnability" (how easy it is to download, build, and run the project - an important measure of reusability). For BOSC 2026, we built bosc-pre-review, an agentic skill that assessed six review criteria, and Runabilly, which builds and tests each project in a disposable Docker container for safety. The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts. After the review period, we surveyed the reviewers to determine how useful they found the pre-review. Most of those who responded said they found it useful, but they preferred to check the AI's conclusions against their own, rather than accepting the AI results unquestioningly.

补充信息

↑