发表机构
University College Dublin(都柏林大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以语言模型智能体为唯一研究者开展长时神经架构设计案例,经三阶段约100次实验改进了视觉Transformer,揭示其生产力阶段结构等四项发现,提出自主研究的未来方向。
AI 中文摘要
我们研究当单个通用大语言模型作为唯一研究者处理长时神经架构设计问题时会发生什么。该智能体接收科学问题、初始假设与动机、计算预算及研究支撑条件(源代码与实验管理、实验跟踪、文献访问、持久记忆),随后在一段较长时间内自主提出、实现、评估并记录实验。该研究分为三个由人工指定过渡阶段,逐步扩展智能体的工具范围或问题规模。在约100个连续实验中,该智能体将一个非标准视觉Transformer从弱基线模型改进为小型基准测试上更强、更高效的模型,以及ImageNet-1K上可用但低于当前最优(SOTA)的模型,同时生成了完整的行为轨迹。我们报告四项发现:(i)生产力呈现清晰的阶段结构:早期快速提升、数十个假设后陷入饱和瓶颈,以及通过扩展动作空间而非改变底层模型实现的瓶颈突破;(ii)早期的一个假设对准确率提升贡献更大,后续改进呈长尾分布;(iii)对贪心式增量假设的偏好主要由工作流程导致:“提交或丢弃”评估规则等价于贪心爬山法,其余偏好源于大胆尝试失败后的风险规避及对熟悉文献的锚定;(iv)该智能体独立重新发现已有结果,并在纯通道注意力这一陌生领域推翻了一项标准设计选择。我们得出结论,本研究中工作流程设计的影响至少与智能体能力相当,并提出多样化搜索、预算内“登月”假设、显式分支及适配领域的重新验证作为未来自主研究的可检验方向。
英文摘要
We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget, and research affordances (source and experiment management, experiment tracking, literature access, and persistent memory), then autonomously proposes, implements, evaluates, and records experiments over an extended period. The study comprises three phases, separated by human-declared transitions, that progressively expand the agent's tool surface or problem scale. Across approximately 100 sequential experiments, the agent improves a non-standard Vision Transformer from a weak baseline to a stronger, efficient model on small benchmarks and a usable but sub-SOTA model on ImageNet-1K, while producing a dense behavioural trace. We report four findings.(i)Productivity exhibits a clear phase structure: rapid early gains, a multi-dozen-hypothesis saturation wall, and recovery, with recovery triggered by expanding the action surface rather than changing the underlying model.(ii)A single early hypothesis contributes more to accuracy gain, with later improvements long-tailed.(iii)The preference for greedy, incremental hypotheses is largely workflow-induced: a commit-or-discard evaluation rule is isomorphic to greedy hill-climbing; the remainder reflects risk aversion after bold failures and anchoring on familiar literature. (iv)The agent independently rediscovers established results and, in the unfamiliar regime of pure channel attention, overturns a standard design choice. We conclude that workflow design was at least as influential as agent capability in this study and propose diversified search, budgeted moonshot hypotheses, explicit forks, and regime-aware re-validation as testable directions for future autonomous research.
CommentsThis work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible