面向LLM智能体基于强化学习的后训练的显式轨迹多样性
Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents
- Microsoft(微软)
- tuyoogame(拓优游戏)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出显式轨迹多样性方法TJPO,通过用户指定的轨迹描述符定义并优化LLM智能体后训练中的行为多样性,在Sokoban和ALFWorld上提升任务性能并保持有效性。
AI中文摘要:
LLM智能体对于同一任务往往存在多个高质量解决方案,这些方案在推理结构、工具使用模式或交互轨迹上有所差异。然而,现有LLM后训练中的多样性概念大多是隐式的,源于一般的随机性和正则化机制,而非明确针对任务相关的行为变异。虽然这种隐式多样性可能有用,但它并未直接指定对于给定任务应鼓励哪些形式的行为变异。在本工作中,我们研究LLM基于强化学习的后训练中的显式轨迹多样性。我们的关键思想是通过用户指定的、任务特定的轨迹描述符来定义多样性,这些描述符将每条采样轨迹映射为可解释的行为表示,然后通过描述符矩阵上的集合级泛函来度量多样性。基于这一表述,我们提出了轨迹引导的联合策略优化(TJPO),这是一个单策略框架,通过轨迹级学习信号在采样轨迹组上优化显式多样性,避免了基于群体的策略训练需求,并将其实例化到基于组的策略优化中。这种设计使多样性目标既可解释又可控。在Sokoban和ALFWorld上的实验表明,TJPO在保持竞争性任务性能的同时,提高了任务特定的轨迹多样性。描述符和轨迹分析表明,学习到的变异遵循指定的行为维度,并包含不同的成功策略。额外的实验结果提示,显式塑造轨迹多样性可以帮助LLM智能体满足用户需求,并在任务条件变化时保持有效性。
英文摘要:
LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation. While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task. In this work, we study explicit trajectory diversity in RL-based post-training for LLMs. Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretable behavioral representation, and then measure diversity as a set-level functional over the resulting descriptor matrix. Building on this formulation, we introduce Trajectory-guided Joint Policy Optimization(TJPO), a single-policy framework that optimizes explicit diversity over sampled trajectory groups, avoiding the need for population-based policy training, and instantiate it within group-based policy optimization through trajectory-level learning signals. This design makes the diversity objective both interpretable and controllable. Experiments on Sokoban and ALFWorld show that TJPO improves task-specific trajectory diversity while maintaining competitive task performance. Descriptor and trajectory analyses show that the learned variation follows the specified behavioral dimensions and includes distinct successful strategies. Extra experiment results suggest that explicitly shaping trajectory diversity can help LLM agents satisfy user requirements and remain effective when task conditions change.