发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器学习工程长期决策难题,提出套娃智能体框架,将问题解决分解为协调层次结构,高级协调器发指令,低级子智能体执行,开发有效训练范式,实验证明该方法有效且可扩展,能提升性能。
AI 中文摘要
机器学习工程(MLE)任务需要在昂贵且由反馈驱动的环境交互下,通过迭代的解决方案调试和优化进行长期决策。开发和训练一个整体式智能体具有根本挑战性,因为它必须同时管理极长且有噪声的上下文,探索广阔的解决方案空间,并在有限的模型容量和计算预算下保持有效。为应对这些挑战,我们提出了套娃智能体,这是一个用于复杂长期任务的统一分层智能体框架。套娃智能体将智能体问题解决分解为决策和执行的协调层次结构:一个高级协调器维护紧凑的长期探索状态并发出战略指令,而低级子智能体通过标准化工具接口介导的直接环境交互执行具体的解决方案尝试。这种设计将战略探索与昂贵的执行解耦,大大减轻了长上下文推理的负担,并实现了高效的迭代优化。我们进一步为套娃智能体开发了一种有效的训练范式。在具有不同模型类型和规模的广泛MLE任务上的实验结果表明,套娃智能体是用于长期MLE任务和复杂智能体问题解决的有效且可扩展的范式。值得注意的是,套娃智能体使Qwen3 - 4B - Instruct达到与o4 - mini相当的协调器性能。将套娃智能体应用于Qwen3 - 30B - Coder最多可带来36.7%的相对性能提升。
英文摘要
Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment interactions. Developing and training a monolithic agent for such tasks is fundamentally challenging, as it must simultaneously manage extremely long and noisy contexts, explore vast solution spaces, and remain effective under limited model capacity and computational budgets. To address these challenges, we propose Matryoshka Agent, a unified hierarchical agent framework for complex long-horizon tasks. Matryoshka Agent decomposes agentic problem solving into a coordinated hierarchy of decision making and execution: a high-level Orchestrator maintains compact, long-horizon exploration states and issues strategic instructions, while lower-level Sub-Agents execute concrete solution attempts through direct environment interaction, mediated by standardized Tool interface. This design decouples strategic exploration from costly execution, substantially reducing the burden of long-context reasoning and enabling efficient iterative refinement. We further develop an efficient training paradigm for Matryoshka Agent. Experimental results on a broad range of MLE tasks with diverse model types and scales demonstrate that Matryoshka Agent is an effective and scalable paradigm for long-horizon MLE tasks and complex agentic problem solving. Notably, Matryoshka Agent enables Qwen3-4B-Instruct to reach Orchestrator performance comparable to o4-mini. Applying Matryoshka Agent to Qwen3-30B-Coder results in at most 36.7% relative performance gain.