arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29641cs.MA

Harness-RL:采用动作-参数解耦的黑盒强化学习,用于中心智能体多智能体管控框架

Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses

  • Peking University(北京大学)
  • National Engineering Research Center of Software Engineering, Peking University(软件工程国家工程研究中心,北京大学)
  • School of Computer Science, Peking University(北京大学计算机学院)
  • Key Laboratory of High Confidence Software Technologies, Ministry of Education(教育部高可信软件技术重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Xinke Jiang, Zhixin Zhang, Zhibang Yang, Jiaran Gao, Rihong Qiu, Shijin Chen, Xu Chu, Junfeng Zhao, Yasha Wang

AI总结:

Harness-RL是解耦动作与参数的黑盒强化学习框架,用于中心智能体多智能体管控框架,在7个基准测试中,使用Qwen2.5模型分别达到42.93和47.79的平均F1分数,验证了CAPO的有效性。

AI中文摘要:

大型语言模型智能体正日益通过多智能体管控框架解决长周期任务,其中中心智能体协调专门的子智能体、工具与环境。训练此类管控框架中的中心策略会面临两大挑战:其一,动作标签是低基数决策,而其参数(args)构成高维条件序列,用共享序列级信号优化两者会产生冲突梯度;其二,动态调度创建的会话存在分支、并行调用与重写上下文,无法忠实简化为单一扁平令牌序列。本文提出Harness-RL,一种结构化强化学习框架,将冲突感知策略优化(CAPO)与接口级黑盒轨迹构建相结合。黑盒组件捕获接口调用记录,构建每个会话的前缀树,并将结果奖励与过程奖励对齐至可训练令牌。CAPO利用前向激活识别与动作及参数令牌相关联的参数分区,再将其策略梯度路由至对应子空间。Harness-RL支持仅中心智能体训练及联合多智能体训练。在7个多跳问答与智能体检索基准测试中,其使用Qwen2.5-1.5B和Qwen2.5-3B分别达到42.93和47.79的平均F1分数, ablation实验验证了CAPO的作用,且在评估设定中倾向于仅中心智能体优化。代码可访问该https URL获取。

英文摘要:

Large language model agents increasingly solve long-horizon tasks through multi-agent harnesses in which a central agent coordinates specialized sub-agents, tools, and environments. Training the central policy in such a harness raises two challenges. First, an action label is a low-cardinality decision, whereas its args form a high-dimensional conditional sequence; optimizing both with a shared sequence-level signal can produce conflicting gradients. Second, dynamic scheduling creates interdependent sessions with branches, parallel calls, and rewritten contexts, which cannot be faithfully reduced to one flat token sequence. We introduce Harness-RL, a structured reinforcement learning framework that combines Conflict-Aware Policy Optimization (CAPO) with interface-level black-box trajectory construction. The black-box component captures Interface Call Records, builds per-session prefix trees, and aligns outcome and process rewards with trainable tokens. CAPO uses forward activations to identify parameter partitions associated with action and args tokens, then routes their policy gradients to the corresponding subspaces. Harness-RL supports both central-only and joint multi-agent training. Across seven multi-hop question answering and agentic retrieval benchmarks, it reaches average F1 scores of 42.93 and 47.79 with Qwen2.5-1.5B and Qwen2.5-3B, respectively, while ablations validate the contribution of CAPO and favor central-only optimization in the evaluated setting. Our code is available at https://github.com/jiangxinke/Harness-RL.

补充信息

↑