发表机构
Meta AI; Columbia University; Tel Aviv University(Meta AI; 哥伦比亚大学; 特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程研究智能体难以学习下一步调查决策的问题,提出层次化元推理架构MIRA,将研究分配与执行分离,并训练生成式演员-评论家MIRA-AC,无需策略训练即可提升推理效率与计算分配,实现高效长时程强化学习。
AI 中文摘要
长时程研究智能体必须在证据积累的过程中决定如何调查以及下一步调查什么。这一决策难以学习,因为此类决策在长执行轨迹中稀疏出现,且其后果可能在数次调查之后才显现。我们提出了迭代研究智能体的元推理(MIRA),一种将研究分配与执行分离的层次化架构。外层循环的元推理器从持久研究记录中整理上下文,然后为下一次调查编写工作指令或终止回合。一个全新的内层循环执行器负责执行每项工作指令,使执行成为元推理动作之间转换的一部分。无需策略训练,MIRA即可改进长时程推理,并在定理证明和开放式神经架构研究中更有效地分配额外计算资源。其决策边界还为信用分配提供了自然单元。在每个边界处,我们训练一个生成式评论家,从部分状态预测预期剩余回报,其性能优于基于词元级别的替代方案。跨环境预训练改善了预测和适应能力,为评估部分进展提供了可迁移的先验。我们利用该先验初始化MIRA-AC,一种生成式演员-评论家模型,联合训练以预测剩余回报并选择下一步调查,无需单独的评论家模型。MIRA-AC将策略优化集中于元推理决策,实现了高效的长时程强化学习,而无需直接优化其启动的更长执行轨迹。在模型自身的代理爬山信号上训练MIRA-AC,提升了四个自动研究环境中的金标准性能;演员通过跨环境价值初始化实现迁移。综合这些结果表明,元推理可以作为显式策略来学习,以指导长时程自主研究。
英文摘要
Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model's own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.