GeoRoute:面向交通未来帧预测的几何感知混合推理方法
GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction
浏览论文内容
中文总结 AI 辅助
GeoRoute是一种无需训练的几何感知混合推理框架,通过多帧深度分层渲染器和视角条件选择预测器,在AI City Challenge Track 5基准上实现了具有竞争力的交通未来帧预测性能,提升了静态几何稳定性。
中文摘要 AI 辅助
长时未来帧预测对自动驾驶、交通监控和智能交通系统至关重要,但因时间重影、几何漂移和物体运动不一致而颇具挑战性。近期的潜在视频扩散模型已实现出色的视觉质量,但将其直接应用于结构化交通场景时,往往会在长时预测中出现几何不稳定和时间连贯性下降的问题。本文提出一种无需训练的推理框架,通过多帧时间上下文和视角条件路由,在预训练视频预测中稳定可靠的静态结构。针对前置摄像头视频,该方法采用多帧深度分层渲染器优化生成的未来帧,该渲染器从观测历史帧中投影静态几何,同时保留生成基础模型的动态区域;针对异构交通视角,冻结的视觉语言模型从观测片段中推断出粗略的相机组,并选择专门的基于运动的预测器。该框架无需对底层视频模型进行重新训练或微调,可直接应用于预训练生成器。我们在AI City Challenge Track 5基准上验证了该框架,最终系统在排名靠前的团队中取得了具有竞争力的性能。这些结果表明,几何感知的推理时优化和视角条件混合推理可在不改变原始模型架构的情况下,提升静态几何稳定性和低级结构保真度。
英文摘要
Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.
发表机构
- Vietnamese-German University(越南-德国大学)
- University of Science, Ho Chi Minh City(胡志明市科学大学)
- Ho Chi Minh City University of Technology(胡志明市技术大学)
- University of Information Technology, Ho Chi Minh City(胡志明市信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。