发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出回合级多尺度密度比估计(tlm-DRE),通过回合级加权和非对称令牌级训练,提升LLM智能体在多回合任务中的对齐性能,并在多个基准上表现优于传统方法。
AI 中文摘要
随着大语言模型(LLM)的快速发展,由LLM增强的智能体系统在处理复杂任务方面展现出巨大潜力,尤其是涉及多步思考或与工具交互的任务。为了以精心设计的智能体范式应用LLM技术,需要在多种智能体场景中对LLM进行后训练以获得更好的性能。在各种后训练技术中,PPO、DPO、DIL和GRPO等对齐方法变得流行,因为许多论文表明,通过惩罚负样本同时保持可接受的训练复杂度,这些方法对模型性能有显著的积极影响。然而,大多数对齐方法处理的是简单的单回合任务,对于复杂的多回合任务仍有改进空间。我们提出了回合级多尺度密度比估计(tlm-DRE),该方法为相应的回合分配不同的权重,并基于多回合任务中的正负空间差距提出非对称的令牌级训练。在广泛的智能体基准上的实验结果表明,与传统对齐方法相比,所提出的方法具有竞争力。所提出的训练方法使LLM能够在领域内和领域外的条件下稳健地执行多回合推理任务。
英文摘要
With the rapid development of Large language model (LLM), agent systems enhanced by LLMs show huge potential in being able to deal with complex tasks, especially involving multi-step thinking or interaction with tools. For applying LLM techniques with a well-designed agent paradigm, post-training of LLM in multiple agent scenarios is necessary to achieve better performance. Among the variable post-training techniques, alignment methods such as PPO, DPO, DIL, and GRPO become popular because many papers show a significant positive impact on the model's performance by punishing negative samples while keeping acceptable training complexity. However, most alignment methods address simple single-turn tasks, and there remains room for improvement for complex multi-turn tasks. We propose Turn-level Multiscale Density Ratio Estimation (tlm-DRE), which assigns different weights on corresponding turns and proposes asymmetric token-level training based on the positive-negative space gaps across multiple turns of tasks. The results of the experiment on a wide range of agent benchmarks show that the proposed method performs competitively compared to traditional alignment methods. The proposed training method enables LLMs to perform robustly in multi-turn reasoning tasks with both in-domain and out-of-domain conditions.
Comments15 pages, 9 figures, 3 tables