通过V-JEPA将语义与字幕生成解耦的仿真到真实交通场景理解
Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA
- Ho Chi Minh City University of Technology and Engineering (HCMUTE)(胡志明市技术工程大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对仿真到真实域迁移下的交通场景理解,提出解耦语义与字幕生成的框架,利用V-JEPA和结构化细化提升VQA与描述生成,在2026挑战赛Track 2中获第一。
AI中文摘要:
AI City Challenge 2026的Track 2要求在具有挑战性的合成到真实域迁移下,同时进行视觉问答(VQA)和交通事件描述生成。现有的视觉-语言方法常常将语义理解与语言生成纠缠在一起,使其容易产生幻觉,并在事件各阶段出现推理不一致。在本工作中,我们提出了一种解耦的语义理解框架,该框架首先将预定义的交通问题解析为结构化的语义事实,随后利用这些事实来指导字幕生成。一个冻结的V-JEPA编码器提取预测性场景表示,而一个轻量级的基于Llama的预测器为VQA查询生成答案。为了提高可靠性,我们引入了一种无需训练的结构化细化机制,该机制利用统计先验、问题间关系和时间事件一致性来纠正预测错误。随后,将细化后的语义事实提供给Qwen3-VL-8B,为每个交通事件生成行人和车辆描述。在官方2026 AI City Challenge Track 2基准上的实验结果表明,所提出的方法达到了87.09%的VQA准确率和60.0853的总体S2分数,在所有参赛队伍中排名第一。这些结果证明,预测性世界表示与结构化语义细化相结合,能够实现更准确、更可靠的交通理解,从而产生更高质量的语言生成。
英文摘要:
Track 2 of the AI City Challenge 2026 requires both visual question answering (VQA) and traffic event description generation under a challenging synthetic-to real domain shift. Existing vision-language approaches often entangle semantic understanding with language generation, making them susceptible to hallucination and inconsistent reasoning across event phases. In this work, we propose a decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation. A frozen V-JEPA encoder extracts predictive scene representations, while a lightweight Llama-based predictor produces answers for VQA queries. To improve reliability, we introduce a training-free structured refinement mechanism that exploits statistical priors, inter-question relationships, and temporal event consistency to correct prediction errors. The refined semantic facts are then provided to Qwen3-VL-8B to generate pedestrian and vehicle descriptions for each traffic event. Experimental results on the official 2026 AI City Challenge Track 2 benchmark show that the proposed method achieves 87.09% VQA accuracy and an overall S2 score of 60.0853, ranking first among all participating teams. These results demonstrate that predictive world representations combined with structured semantic refinement enable more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.