HarnessBandit:多框架智能体强化学习的联合可学习性-可迁移性调度
HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对多框架智能体强化学习中的调度问题,提出HarnessBandit在线调度器,融合可学习性与可迁移性信号,在ClawGym训练并在两个基准上超越混合批次训练。
中文摘要 AI 辅助
语言模型智能体越来越多地通过不同的框架进行部署,这些框架在系统提示、工具模式、控制循环和轨迹格式上各不相同。同一模型在这些接口上的表现可能参差不齐,因此对框架变化的鲁棒性成为一个重要目标。一种自然的方法是通过多个框架训练一个共享策略,但这引入了调度问题:每个训练步骤应优先选择当前能提供有用学习信号的框架,同时产生的更新也要有益于其他框架。我们开发了HarnessBandit,一个在线调度器,每个优化器步骤选择一个框架。在组相对策略优化(GRPO)更新后,它观察可学习性——批次上的平均绝对优势——以及可迁移性——当前框架的低维梯度草图与其余框架的指数移动平均之间的余弦相似度。这两个信号在池化滑动窗口最小-最大归一化后融合,并通过访问相关奖励和显式探索下限进行采样。我们在ClawGym上使用六个框架训练Qwen3.5-2B,并在PinchBench(保留任务,分布内OpenClaw)和ClawEval(保留任务和框架)上进行评估。HarnessBandit在两个基准上都优于混合批次多框架训练,而训练诊断表明可学习性和可迁移性提供了不同且不断演变的信号。
英文摘要
Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory formats. The same model can perform unevenly across these interfaces, making robustness to harness variation an important objective. A natural approach is to train a shared policy through multiple harnesses, but doing so introduces a scheduling problem: each training step should favor a harness that currently provides a useful learning signal while also producing an update that benefits the other harnesses. We develop HarnessBandit, an online scheduler that selects one harness per optimizer step. After a group-relative policy optimization (GRPO) update, it observes learnability -- the mean absolute advantage on the batch -- and transferability -- the cosine between a low-dimensional gradient sketch of the current harness and exponential moving averages of the remaining harnesses. The two signals are fused after pooled sliding-window min-max normalization and sampled with a visit-dependent bonus and an explicit exploration floor. We train Qwen3.5-2B across six harnesses on ClawGym and evaluate on PinchBench (held-out tasks, in-distribution OpenClaw) and ClawEval (held-out tasks and harness). HarnessBandit improves over mixed-batch multi-harness training on both benchmarks, while training diagnostics indicate that learnability and transferability provide distinct, evolving signals.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- Alibaba Cloud(阿里云)
机构由 AI 辅助整理,请以论文原文为准。