AI 中文总结
本文提出主动子结构感知策略优化(ASPO),一种强化学习框架,通过自适应查询选择和子结构级奖励,在查询级搜索多智能体系统架构,在六个基准上均超越十二个基线。
AI 中文摘要
大型语言模型(LLMs)使得多智能体系统(MAS)能够处理复杂任务,但手动设计智能体角色、提示词和通信结构需要大量的专业知识和努力。这促使我们学习从执行奖励中构建特定于查询的MAS的策略。现有方法通常通过反复遍历一组固定的训练查询并在工作流级别分配奖励来训练这些策略。然而,这忽略了查询在演化学习潜力上的差异,并掩盖了哪些子结构能提高解决方案质量。在本文中,我们提出了主动子结构感知策略优化(ASPO),一个用于查询级MAS搜索的强化学习框架。ASPO引入了自适应查询选择机制(AQSM),将训练集中在策略能力边界上的查询上:即那些能够解决但尚不可靠的查询。一个补充的发现机制扩展了对困难查询的架构探索,帮助区分探索不足与操作者能力限制。除了查询选择,ASPO还引入了子结构级奖励,用于衡量每个动作的后代子图内的输出质量增益。这些奖励指导近端策略优化,以强化有用的架构改进并抑制冗余或有害的计算。这些机制共同优先考虑可学习的查询,并为学习有效的MAS提供细粒度的反馈。在涵盖数学推理、通用问答和代码生成的六个基准测试中,ASPO在每一个基准测试上均排名第一,超过了十二个基线。
英文摘要
LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically train these policies by repeatedly traversing a fixed set of training queries and assigning rewards at the workflow level. However, this overlooks differences in queries' evolving learning potential and obscures which substructures improve solution quality. In this paper, we propose Active Substructure-aware Policy Optimization (ASPO), a RL framework for query-level MAS search. ASPO introduces an Adaptive Query-Selection Mechanism (AQSM) that focuses training on queries at the policy's competence boundary: those it can solve but not yet reliably. A complementary discovery mechanism widens architectural exploration for hard queries, helping distinguish insufficient exploration from operator capability limits. Beyond query selection, ASPO introduces substructure-level rewards that measure output-quality gains within each action's descendant subgraph. These rewards guide proximal policy optimization to reinforce useful architectural refinements and discourage redundant or harmful computation. Together, these mechanisms prioritize learnable queries and provide fine-grained feedback for learning effective MASs. Across six benchmarks spanning mathematical reasoning, general question answering, and code generation, ASPO ranks first on every benchmark against twelve baselines.