智能体多轮推理:一种公平性方法
Agentic Multi-Turn Reasoning: A Fairness Approach
浏览论文内容
中文总结 AI 辅助
针对智能体多轮推理中的长时程信用分配和数据不平衡问题,提出公平多级偏好优化框架(Fair-MPO),理论分析与实验证明其达到最先进性能。
中文摘要 AI 辅助
大型语言模型(LLMs)的最新进展使得智能体系统能够通过多轮规划、工具使用、验证和记忆更新来解决复杂任务。然而,由于两个基本挑战,学习智能体系统仍然困难,即(1)长时程信用分配,其中监督仅在最终结果处可用,以及(2)不平衡的数据分布,其中占主导地位的数据模式使优化产生偏差,并削弱对罕见但信息丰富的推理行为的适应。在本文中,我们提出了公平多级偏好优化(Fair-MPO或Φ-MPO),一种用于智能体学习的新型偏好优化框架。我们首先表明,多级偏好优化为长时程推理提供了一个有原则且计算效率更高的框架。然后,我们引入了一个公平多级目标,以解决智能体学习中的不平衡问题。我们提供了全面的理论分析,证明我们的方法同时解决了长时程推理和数据不平衡问题。我们在智能体推理基准上的实验表明,我们的方法达到了最先进的(SOTA)性能。
英文摘要
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $Φ$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.