Atria Dawn:代理式超级智能的黎明
Atria Dawn: The Dawn of Agentic Superintelligence
浏览论文内容
中文总结 AI 辅助
本文介绍Atria Dawn Preview,一个通过可验证经验管道训练的代理式语言模型,在16个基准中表现领先,并通过人机协作案例分析揭示从任务执行向项目伙伴关系的转变。
中文摘要 AI 辅助
随着AI代理成为其继任者开发过程中的参与者,它们既重塑了智能的生产方式,也重塑了人类研究者的角色。我们推出Atria Dawn Preview,一个为科学研究和工程工作流设计的基础代理式语言模型,其目标是拓展代理在现实世界中生产力的前沿。该模型通过可验证经验管道(Verifiable Experience Pipeline)进行训练,该管道将工具介导的交互连接到可执行环境和外部验证的结果。在涵盖现实世界研究、工程和数字工作的16个基准测试中,Atria Dawn Preview与前沿代理相比具有竞争力,并在其中五个基准上取得了最高报告分数。除了独立性能之外,我们将该模型背后的实际研发过程作为人机协作的案例研究,分析了来自56名参与者的769条任务记录以及代理日志。当被要求在可比条件下评估已完成任务时,参与者将约三分之一的已完成AI辅助任务评为在没有AI的情况下不可行。更引人注目的是,代理经常提出方法并实施修订,而人类保留大多数最终决策,并通过判断和反馈引导探索。这些观察表明,从任务级执行向项目级伙伴关系的转变,人类的努力集中在什么值得追求以及证据应如何指导研究上。因此,迈向更自主的AI研究的进展必须同时提升发现能力和有意义的人类监督能力,在持续开发的风险和方向上保留负责任的人类权威。
英文摘要
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.