arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI智能体时代需要新的科学范式以维持可信科学

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

Belinda Mo

arXiv 2607.26064首次发表:更新:

发表机构

Long Horizon Research(长 horizon 研究机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI自主研究智能体带来的科学产出验证缺口扩大问题,该文提出需进化科学验证基础设施的标准,以应对无人能核查结果等风险,维持科学信任。

AI 中文摘要

AI系统正成为自主研究智能体,能生成假说、设计实验并以超出人类监督的规模产出发现。从机器学习(ML)会议投稿量增加可见,科学产出与我们核查其的能力间的验证缺口已在扩大,而考虑到人类与智能体的不对称性,自主智能体将使这一缺口大幅加剧。我们认为科学必须像过往应对同行评审那样,进化其验证基础设施。不过,过往的调整假设了可被质询和制裁的人类贡献者,而AI智能体打破了这一假设。我们提出适配后的验证基础设施的标准,强调默认可观测的工作流程、可扩展的验证及清晰的归因。我们认为若不调整,使用智能体的ML及任何科学领域将面临危险失败:无人能核查的实验结果、为指标而非理解优化、侵蚀科学信任的责任真空。

英文摘要

AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it has before with peer review. However, while historical adaptations assumed human contributors who could be questioned and sanctioned, AI agents break this assumption. We propose criteria for an adapted verification infrastructure that emphasizes observable-by-default workflows, scalable verification, and clear attribution. We argue that without adaptation, ML and any scientific domain using agents face dangerous failures: experimental results that no person can verify, optimization for metrics over understanding, and accountability vacuums that erode scientific trust.

CommentsAccepted at ICML 2026, Position Paper Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑