arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14399cs.AIcs.CL

MOSCOPT:面向LLM智能体的技能混合集体优化

MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

Zhenyu Zhang, Jiudong Yang

首次发表
浏览论文内容

中文总结 AI 辅助

提出MOSCOPT,一种无参数算法,通过联合优化技能池和门控技能,利用EditAdam三阶段交错更新,在多个基准上超越基线,实现LLM智能体的集体优化。

中文摘要 AI 辅助

自然语言提示和技能构成了基于LLM的智能体的战略支柱。近期在提示和技能优化方面的进展取得了显著成果,然而所有现有方法都优化一个单一文本模板——忽略了多种互补策略之间的协同效应。我们提出MOSCOPT,一种文本原生、无参数的算法,它联合优化一个包含N个技能的技能池和一个门控技能G,后者在每一步动态选择K个技能。为了有效优化这些技能,我们构建了带有内部维护的双状态的EditAdam。通过使用EditAdam进行的三阶段交错更新,系统在无需梯度或参数调整的情况下单调改进。在5个基准和3个目标LLM上的广泛实验和详细消融研究表明,MOSCOPT始终优于所有基线,并证实了具有选择性激活的技能混合架构和具有三阶段交错性的集体进化对其卓越性能至关重要。代码已在此https URL发布。

英文摘要

Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a \emph{single} text template---missing the synergy among multiple complementary strategies. We propose MOSCOPT, a text-native, parameter-free algorithm that jointly optimizes a pool of $N$ skills and a gating skill $G$ that dynamically selects $K$ skills per step. To effectively optimize the skills, we build the EditAdam with internally maintained dual states. Through the three-phase interleaved updates with EditAdam, the system monotonically improves without gradient or parameter tuning. Extensive experiments and detailed ablations across 5 benchmarks and 3 target LLMs demonstrate that MOSCOPT consistently outperforms all baselines, and confirm that both the mixture-of-skills architecture with selective activation and the collective evolution with three-phase interleaving are essential to its superior performance. Code is released https://github.com/zhangzhenyu13/SummerClaw/tree/master/summerclaw/agent_trainer/algorithms/moscopt.

发表机构

  • ML Center, Coupang(Coupang机器学习中心)
  • AI Center, Futu AI(富途AI人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑