发表机构
Yonsei University; DGIST(延世大学; 大邱庆北科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于语言的轨迹预测中坐标动态建模不足,提出MoRE框架,通过强化学习融合五个数值预测器的坐标级知识,在ETH-UCY等数据集上显著降低ADE和FDE,并改善碰撞率与行人间距。
AI 中文摘要
基于语言的轨迹预测器将坐标表示为离散标记,并学习辅助任务,如目的地和群体推理。这种表述使模型能够捕获超越坐标动态的行为意图和社会背景。然而,标记级目标仅对连续坐标空间动态提供间接指导。为解决这一局限性,我们引入了MoRE(Mixture of Reward Experts),一个通过强化学习将数值预测先验转移到预训练的基于语言的预测器中的细化框架。五个冻结的数值预测器提供互补的坐标级运动和交互知识。它们的预测被转换为专家奖励,并通过不确定性加权共识组合,该共识惩罚分歧。一个地面真值奖励将预测锚定到目标轨迹。为了将细化集中在困难案例上,MoRE使用按预测熵排序的前1%训练样本细化策略。专家预测在PPO训练前计算一次并缓存,因此在策略更新或推理期间不运行专家。这样,MoRE结合了基于语言的预测器的上下文建模与数值专家的坐标级反馈。在ETH-UCY上,MoRE将ADE从0.22米降至0.20米,FDE从0.32米降至0.29米。相对于基础策略,ADE在SDD上降低17.9%,在NBA上降低12.7%。在ETH-UCY上,MoRE还减少了碰撞率,并更好地匹配地面真值行人间距,而不增加测量的推理内存或延迟。项目页面可从此https URL访问。
英文摘要
Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning. This formulation enables the model to capture behavioral intent and social context beyond coordinate dynamics alone. However, token-level objectives provide only indirect guidance for continuous coordinate-space dynamics. To address this limitation, we introduce MoRE (Mixture of Reward Experts), a refinement framework that transfers numerical forecasting priors into a pretrained language-based predictor through reinforcement learning. Five frozen numerical predictors provide complementary coordinate-level knowledge of motion and interactions. Their predictions are converted into expert rewards and combined through an uncertainty-weighted consensus that penalizes disagreement. A ground-truth reward anchors the prediction to the target trajectory. To focus refinement on difficult cases, MoRE refines the policy using the top 1% of training samples ranked by predictive entropy. Expert predictions are computed once and cached before PPO training, so the experts are not run during policy updates or inference. In this way, MoRE combines the contextual modeling of the language-based predictor with coordinate-level feedback from numerical experts. On ETH-UCY, MoRE reduces ADE from 0.22 to 0.20 m and FDE from 0.32 to 0.29 m. Relative to the base policy, ADE decreases by 17.9% on SDD and 12.7% on NBA. On ETH-UCY, MoRE also reduces collision rates and better matches ground-truth pedestrian spacing, without increasing measured inference memory or latency. The project page is available at https://jungyu0413.github.io/MoRE/.
Comments35 pages, 15 figures. Project page: https://jungyu0413.github.io/MoRE/