arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

直接多样性优化:偏好后训练中多样化成功轨迹的生成

Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

Junwon Ko, Dong-Jae Lee, Minchan Kwon, Sunghyun Baek, Junmo Kim

arXiv 2609.10052首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出直接多样性优化(DDO)方法,结合分歧树收集与参考相对目标几率目标,提升LLM在顺序决策任务中成功策略的覆盖多样性,实验显示优于现有后训练方法。

AI 中文摘要

用于顺序决策任务的LLM智能体通常在轨迹级结果标签上进行后训练,但此类标签对保留来自同一决策状态的多个成功分支提供的监督很少。我们将此问题研究为成功策略覆盖:在固定 rollout 预算下,模型实现不同成功策略的广度。我们提出了直接多样性优化(DDO),一种离线后训练方法,将分歧树收集(DTC)与参考相对目标几率目标(RTO)相结合。DTC构建以共享决策状态为根的状态对齐分支集,RTO训练模型以匹配成功备选方案上的参考相对目标。在BabyAI、BabaIsAI和WebShop上,DDO在比较的后训练方法中实现了最强的任务成功率和成功策略覆盖。它还在局部动作替换后实现了最高的恢复率,并且比仅成功模仿和解码时多样化控制具有更高的任务成功率和覆盖率。

英文摘要

LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverage: how broadly a model realizes distinct successful strategies under a fixed rollout budget. We present Direct Diversity Optimization (DDO), an offline post-training method that combines Divergence-Tree Collection (DTC) with the Reference-Relative Target-Odds Objective (RTO). DTC constructs state-aligned branch sets rooted at shared decision states, and RTO trains the model to match reference-relative targets over successful alternatives. DDO achieves the strongest task success and successful strategy coverage among the compared post-training methods across BabyAI, BabaIsAI, and WebShop. It also achieves the highest recovery rate after local action replacement and higher task success and coverage than successful-only imitation and decoding-time diversification controls.

CommentsAccepted to EMNLP 2026 Main Conference. 19 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑