arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10601cs.AI

Agentic-DPO:从专家轨迹上的模仿到智能体策略优化

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories

Yixiong Chen, Alan Yuille

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型智能体基于专家轨迹训练只学动作序列、难应对错误的问题,提出Agentic-DPO方法,通过转化专家轨迹为状态条件偏好监督,结合策略保持增强,实现低成本智能体策略优化,实验验证其有效性。

中文摘要 AI 辅助

大语言模型智能体通常通过监督微调(SFT)从专家轨迹进行训练,将多轮智能体行为视为普通文本模仿。这种方法简单低成本,但只学习模仿专家动作序列,而非训练智能体在每个状态针对可能的错误选择正确动作。现有缓解此问题的方法包括偏好学习或强化学习,但通常需要高成本的环境展开和奖励模型。我们提出Agentic-DPO,一种轻量级离线智能体策略优化方法,将专家轨迹转化为状态条件偏好监督。在每个专家动作状态,Agentic-DPO从当前状态采样一个单步动作,将可能的错误动作视为负样本,并使用DPO风格的偏好目标与专家动作进行对比。为避免在偏好学习中混合策略和模式,我们引入策略保持增强(PPA),在保持专家策略不变的同时,在多个模式下生成相同的潜在轨迹。Agentic-DPO无需在线环境展开、奖励模型或全轨迹学生探索。我们在StableToolBench、tau-bench零售和Mind2Web上进行实验,Agentic-DPO在不同模型规模下持续改进智能体,超越模仿效果。特别是,对于9B模型,它将tau-bench准确率从21.7%(SFT)提高到41.4%,在相同骨干网络下仅通过步级展开且在梯度步骤中无环境交互就匹配了在线GRPO。结果表明,当专家轨迹从示范转换为状态级动作偏好时,可支持低成本的智能体策略优化。Agentic-DPO的代码在这个https URL发布。

英文摘要

Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitation. This recipe is simple and low-cost, but it only learns to imitate the sequence of expert actions, rather than training the agent to choose the right action against plausible mistakes at each state. Existing methods to mitigate this problem include preference learning or reinforcement learning, but they usually need high-cost environment rollouts and reward models. We propose Agentic-DPO, a lightweight offline agent policy optimization method that turns expert trajectories into state-conditioned preference supervision. At each expert action state, Agentic-DPO samples a one-step action from the current state, treats plausible wrong actions as negatives, and contrasts them with the expert action using a DPO-style preference objective. To avoid mixing both policy and schema in preference learning, we introduce Policy-Preserving Augmentation (PPA), which renders the same latent trajectory under multiple schemas while keeping the expert policy fixed. Agentic-DPO requires no online environment rollout, reward model, or full-trajectory student exploration. We conduct experiments across StableToolBench, tau-bench retail, and Mind2Web, where Agentic-DPO consistently improves agents at different model scales beyond imitation. In particular, it raises tau-bench accuracy from 21.7% (SFT) to 41.4% for a 9B model, matching online GRPO under the same backbone with only step-level rollouts and without environment interaction during gradient steps. The results suggest that expert trajectories can support low-cost agentic policy optimization when converted from demonstrations into state-level action preferences. Code for Agentic-DPO is released at https://github.com/Schuture/Agentic-DPO.

发表机构

  • Johns Hopkins University(约翰·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑