arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoCAR:面向连续轨迹预测的运动码坐标感知自回归

MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting

Yiming Xu, Hao Cheng, Monika Sester

arXiv 2610.06210首次发表:更新:

发表机构

Leibniz University Hannover; University of Twente(莱布尼茨汉诺威大学; 特文特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoCAR提出坐标感知连续潜在空间中的运动码自回归,将轨迹预测转化为下一码预测,实现无需重标记的单阶段架构,在Argoverse基准上达到顶级性能并强迁移。

AI 中文摘要

自回归生成对于语言是自然的,其中预测的标记可以直接重用为下一个预测状态,但轨迹预测缺乏这样干净的标记:运动是连续的、多模态的,并在随预测轨迹演化的局部坐标框架中表达。我们提出MoCAR(运动码坐标感知自回归),一个仅解码器框架,将轨迹预测转化为坐标感知连续潜在空间中的下一个码预测。MoCAR从端点归一化的轨迹段学习一个连续运动码空间,其中每个码联合捕获局部轨迹几何和由该段引起的参考框架转换。历史运动码用作教师强制前缀,未来码在时间、地图、智能体和模式交互下自回归生成,预测码在潜在记忆中持续存在,而解码的端点更新局部场景上下文。这使得无需轨迹空间重新标记、轨迹查询、目标候选或提议-细化流水线即可进行展开。在Argoverse(AV)基准上,MoCAR以简单的单阶段架构实现顶级性能,从AV2到AV1的零样本评估中强迁移,并在转弯密集场景中改进。消融研究确认,学习的连续运动码空间、潜在对齐、弱KL正则化和联合分词器-预测器优化对于稳定的潜在自回归至关重要。

英文摘要

Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that evolve with the predicted trajectory. We present MoCAR (Motion-code Coordinate-aware AutoRegression), a decoder-only framework that casts trajectory forecasting as next-code prediction in a coordinate-aware continuous latent space. MoCAR learns a continuous motion-code space from endpoint-normalized trajectory segments, where each code jointly captures local trajectory geometry and the reference-frame transition induced by that segment. Historical motion codes are used as a teacher-forced prefix, future codes are generated autoregressively under temporal, map, agent, and mode interactions, and predicted codes persist in latent memory while decoded endpoints update the local scene context. This enables rollout without trajectory-space re-tokenization, trajectory queries, goal candidates, or proposal-and-refinement pipelines. On Argoverse (AV) benchmarks, MoCAR achieves top-tier performance with a simple single-stage architecture, transfers strongly from AV2 to AV1 in zero-shot evaluation, and improves on turn-heavy scenarios. Ablations confirm that the learned continuous motion-code space, latent alignment, weak KL regularization, and joint tokenizer-predictor optimization are essential for stable latent autoregression.

CommentsAccepted at NeurIPS 2026. Camera-ready version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑