通过多智能体偏好学习改进LLM协作
Improving LLM Collaboration via Multi-Agent Preference Learning
中文总结 AI 辅助
本文提出多智能体偏好学习框架MAPL,通过去中心化与中心化协作视角及偏好优化,提升LLM协作质量与效率,实验验证其接近固定奖励MARL性能。
中文摘要 AI 辅助
已有若干工作探索了在LLM协作中的多智能体强化学习(MARL)。然而,在实践中构建可靠的奖励是困难的,因为完整且准确的指标往往不可用且难以聚合。偏好学习通过从比较性的人类或AI反馈中学习提供了一种替代方案。然而,其在多智能体系统中的扩展仍未得到充分探索。为填补这一空白,我们从去中心化和中心化协作的角度构建了基于偏好的多智能体系统(MAS)。我们还引入了一个通用的多智能体偏好学习框架(MAPL)来解决这些问题。MAPL允许通过将当前解决方案与由各种智能体生成的去中心化或中心化解决方案进行比较来进行迭代更新。我们使用从人类反馈中学习的奖励模型的MARL(MARLHF)和多智能体直接偏好优化(MADPO)来实例化MAPL。在协作写作、编码、工具使用和旅行规划上的实验表明,MAPL可以提高协作质量和效率,同时接近具有固定、明确奖励的MARL的性能。在MAPL中,MARLHF在大多数任务上通常优于MADPO,但对数据覆盖范围、智能体和比较模型以及底层MARL算法仍然敏感。
英文摘要
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comparative human or AI feedback. Yet, its extension to multi-agent systems remains underexplored. To address this gap, we formulate preference-based multi-agent systems (MAS) from decentralized and centralized collaboration perspectives. We also introduce a general multi-agent preference learning framework (MAPL) to solve these problems. MAPL allows iterative updates by comparing the current solution with decentralized or centralized solutions generated by various agents. We instantiate MAPL using MARL from human feedback (MARLHF) with a learned reward model and multi-agent direct preference optimization (MADPO). Experiments on collaborative writing, coding, tool use, and travel planning show that MAPL can improve collaboration quality and efficiency while approaching the performance of MARL with fixed, well-defined rewards. Within MAPL, MARLHF generally outperforms MADPO on most tasks but remains sensitive to data coverage, agent and comparator models, and the underlying MARL algorithms.