arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeurIPS与ICML 2025立场论文 track 的调查

An Investigation of the NeurIPS and ICML 2025 Position Tracks

Fan Yang, Wenkai Li, Jun Liu

arXiv 2608.16894首次发表:更新:

发表机构

Fujitsu Research; Carnegie Mellon University(富士通研究院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文审计了NeurIPS和ICML 2025立场论文 track 的公开投稿,发现其以改革派批判为主,建议增设方向设定类工作,并提出四项CFP层面干预措施以优化投稿结构。

AI 中文摘要

机器学习(ML) venues 决定了审稿人认可何种类型的研究主张,以及何种形式的证据被视为严谨。NeurIPS和ICML的立场论文 track 是为设定研究议程的工作设立的,因此其早期构成值得审视。本文指出,2025年公开可获取的已评审投稿以改革派批判为主,该 track 应明确征集方向设定类工作,而非仅保留已有的改革派批判内容。我们按照预先设定的标准,对NeurIPS 2025和ICML 2025立场论文 track 的所有可获取投稿进行了审计,并将所得模式与公认的议程转变类ML论文(如AlexNet、Transformer、《人工智能安全的具体问题》等)的参考组进行了比较。四分之三的被审计投稿对现有基准、评估方法或方法论提出了批判;这些论文在我们的“与人工制品耦合”标准上得分很高,但证据深度并不能预测审稿人的评分。参考组与公开已评审投稿池在“人工制品类型”上存在差异:议程转变类论文通常为该领域提供了可构建、测试或批判的新内容,例如测量协议、基准提案、玩具实现、数据集卡片、审计模板或可证伪的实验方案。我们最后提出了四项针对征集提案(CFP)层面的干预措施,旨在扩大投稿组合,同时不取代该 track 已妥善接纳的批判内容。

英文摘要

ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were created for agenda-setting work, making their early composition worth auditing. \textbf{This paper argues that the publicly accessible 2025 reviewed pool is dominated by reformist critique, and that the track should explicitly solicit direction-setting work alongside, not in place of, the reformist critiques it already hosts well.} We audit every accessible submission to the NeurIPS 2025 and ICML 2025 Position Tracks under a pre-specified rubric, and compare the resulting pattern with a reference class of widely recognized agenda-shifting ML papers. Three-quarters of audited submissions critique an existing benchmark, evaluation, or methodology; these papers score highly on our artifact-coupling rubric, but evidentiary depth does not predict reviewer rating. The reference class (AlexNet, the Transformer, Concrete Problems in AI Safety, and others) differs from the accessible reviewed pool in \emph{artifact kind}: agenda-shifting papers typically gave the field something new to build on, test against, or contest, such as a measurement protocol, benchmark proposal, toy implementation, dataset card, audit template, or falsifiable experimental program. We close with four CFP-level interventions aimed at broadening the submission mix without displacing the critiques the track already hosts well.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑