学习激活什么:面向长时程多模态智能体的组合能力分配
Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents
AI总结:
本文提出CoCA框架,通过教师-学生在线策略学习与双层蒸馏,解决长时程多模态智能体的组合能力分配问题,实现低成本、高稳定的能力子集选择。
AI中文摘要:
长时程多模态智能体依赖感知、检索、推理、验证和执行等专门能力。现有设计通常激活固定的能力集或调用预定义的工作流,这会产生大量的计算开销,同时无法适应随阶段变化的能力需求。在本文中,我们研究长时程多模态智能体系统中的组合能力分配问题,其中系统在每个交互阶段选择一个对成本敏感的专业能力子集,这并非易事,因为能力值取决于所选择的子集,而先前的分配会改变后续决策所遇到的状态。我们引入了CoCA,一个在线策略学习框架,它从稀疏的条件比较中恢复出一个可部署的能力子集策略。在学生策略访问的状态上,更强的教师基于当前选择的子集,比较候选能力的边际净值。然后,我们采用条件效用模型将这些比较转化为自回归的能力子集策略,避免了显式枚举。我们进一步引入双层在线策略蒸馏,以解决环境状态之间以及集合构建过程中遇到的局部子集内的分布不匹配问题。最后,轨迹级强化学习将蒸馏后的策略细化,以优化任务成功率、激活成本和分配稳定性。在推理时,分配仅由轻量级学生策略执行,无需教师查询或在线更新。在长时程多模态环境和受控能力需求变化上的实验证明了我们的方法优于最先进的基线方法。
英文摘要:
Long-horizon multimodal agents rely on specialized capabilities for perception, retrieval, reasoning, verification, and execution. Existing designs typically activate a fixed capability set or invoke a predefined workflow, incurring substantial computational overhead while failing to accommodate stage-dependent capability demands. In this paper, we study the \textit{combinatorial capability allocation} problem for long-horizon multimodal agent systems, where the system selects a cost-sensitive subset of specialized capabilities at each interaction stage, which is nontrivial since capability values depend on the selected subset, while previous allocations alter the states encountered by subsequent decisions. We introduce \textsc{CoCA}, an on-policy learning framework that recovers a deployable capability-subset policy from sparse conditional comparisons. On states visited by the student policy, the stronger teacher compares the marginal net values of candidate capabilities, conditioned on the currently selected subset. Then, we adopt a conditional utility model to transform such comparisons into an autoregressive capability-subset policy, avoiding explicit enumeration. We further introduce dual-level on-policy distillation to address distribution mismatch both across environment states and within the partial subsets encountered during set construction. Finally, trajectory-level reinforcement learning refines the distilled policy toward task success, activation cost, and allocation stability. At inference time, allocation is performed solely by the lightweight student policy without teacher queries or online updates. Experiments on long-horizon multimodal environments and controlled capability-demand shifts demonstrate the superiority of our method over the state-of-the-art baseline methods.