arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一轨迹匹配策略优化:多样化T2I生成与VLA泛化

Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization

Zhiyuan Ma, Jiaming Li, Lingzhen Li, Yu Liu, Xuekai Zhu, Dingkang Liang, Kaiyan Zhang, Jianjun Li, Bowen Zhou, Xiang Bai

arXiv 2609.34688首次发表:更新:

发表机构

Huazhong University of Science and Technology; Institute of Information Engineering, Chinese Academy of Sciences; Shanghai Jiao Tong University; Tsinghua University(华中科技大学; 中国科学院信息工程研究所; 上海交通大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对奖励最大化RL导致T2I和VLA策略模式坍缩的问题,提出Uni-TMPO统一后训练框架,通过前向KL匹配轨迹分布,实现更高奖励、多样性和泛化。

AI 中文摘要

奖励最大化的强化学习(RL)被广泛用于对文本到图像(T2I)生成的随机扩散和流策略进行后训练。然而,即使在参考KL或熵正则化下,奖励最大化的RL也会导致策略模式坍缩,将策略缩减为单一的高奖励模式。在T2I中,这会产生相似的图像和奖励黑客行为。当扩展到视觉-语言-动作(VLA)模型时,同样的坍缩会消除替代的成功策略,并削弱任务和场景的泛化能力。为解决这一局限性,我们引入了统一轨迹匹配策略优化(Uni-TMPO),一个用于扩散和流策略的统一RL后训练框架。首先,Uni-TMPO将标准化奖励转换为每个轨迹组内的目标分布,并从轨迹对数概率中推导策略分布。然后,前向Kullback-Leibler优化匹配这两个分布,而不是最大化期望奖励。一个进度条件的从粗到细调度器高效地构建T2I轨迹。在统一框架内,反馈条件采样使用更新的观测来构建VLA轨迹。大量实验表明,Uni-TMPO在T2I奖励和VLA ID成功率上优于最强基线。更重要的是,它实现了最佳的T2I奖励-多样性-效率权衡以及对保留任务和场景的VLA泛化,而真实机器人评估证明了当高奖励目标被阻塞时多种行动策略的价值。

英文摘要

Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-reward mode. In T2I, this produces similar images and reward hacking. When extended to vision-language-action (VLA) models, the same collapse removes alternative successful strategies and weakens task and scene generalization. To address this limitation, we introduce Unified Trajectory Matching Policy Optimization (Uni-TMPO), a unified RL post-training framework for diffusion and flow policies. First, Uni-TMPO converts standardized rewards into a target distribution within each trajectory group and derives the policy distribution from trajectory log probabilities. Then, forward Kullback-Leibler optimization matches the two distributions instead of maximizing expected reward. A progress-conditioned coarse-to-fine scheduler efficiently constructs T2I trajectories. Within the unified framework, feedback-conditioned sampling uses updated observations to construct VLA trajectories. Extensive experiments show that Uni-TMPO achieves higher T2I rewards and VLA ID success rates than the strongest baselines. More importantly, it achieves the best T2I reward-diversity-efficiency trade-off and VLA generalization to held-out tasks and scenes, while real-robot evaluation demonstrates the value of multiple action strategies when the higher-reward target is blocked.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑