发表机构
The Hong Kong University of Science and Technology (Guangzhou); The Hong Kong University of Science and Technology; Jiangnan University; Central China Normal University; CUHK; Peking University; Wuhan University; Fudan University(香港科技大学(广州); 香港科技大学; 江南大学; 华中师范大学; 香港中文大学; 北京大学; 武汉大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对路边交通场景中接地多模态大模型的说-点不匹配问题,提出先枚举后回答(EtA)和枚举一致性策略优化(ECPO),显著提升答案与证据的一致性及接地F1,并验证了跨场景迁移能力。
AI 中文摘要
路边交通推理要求每个自由形式的文本主张都有视觉证据支持。现有的接地多模态大语言模型(MLLMs)经常表现出“说-点不匹配”问题,即文本答案与模型定位的边界框相矛盾。分别对答案和边界框进行评分的评估指标无法惩罚这种失败。我们将这种不匹配追溯到传统的“先答案后接地”分解方式,该方式在枚举任何对象之前就承诺了一个数值答案。为了衡量这一问题,我们构建了RoadSceneVQA-G基准,包含34.7K个问答对,其中每个自由形式答案都与支持它的边界框集合相关联,并提出了答案-接地一致性(AGC)评估套件。为了解决这一问题,我们引入了“先枚举后回答”(EtA)方法,该方法反转生成顺序,使答案-证据一致性成为输出结构的属性,以及枚举一致性策略优化(ECPO),这是一种强化学习阶段,使用多次采样的并集作为无需真实边界框的召回教师。EtA将说-点一致性从26.6%提高到93.7%,接地F1从52.2%提高到73.0%,ECPO在没有逐框监督的情况下进一步将F1提高到75.6%。在gRefCOCO上,相同的框架优于最强对比方法,表明其可迁移到交通场景之外。项目可在该URL获取。
英文摘要
Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and boxes separately leave this failure unpenalized. We trace the mismatch to the conventional answer-then-ground factorization, which commits to a numerical claim before any object is enumerated. To measure it, we build RoadSceneVQA-G, a benchmark of 34.7K question-answer pairs in which every free-form answer is linked to the set of boxes that witnesses it, and we propose the Answer-Grounding Consistency (AGC) evaluation suite. To address it, we introduce Enumerate-then-Answer (EtA), which reverses the generation order so that answer-evidence agreement becomes a property of the output structure, and Enumeration-Consistent Policy Optimization (ECPO), a reinforcement learning stage that uses the union of multiple rollouts as a recall teacher without ground-truth boxes. EtA raises say-point consistency from 26.6\% to 93.7\% and grounding F1 from 52.2\% to 73.0\%, and ECPO further increases F1 to 75.6\% without per-box supervision. On gRefCOCO, the same framework outperforms the strongest compared method, indicating that it transfers beyond traffic scenes. The project is available at \url{https://github.com/GuanRunwei/RoadSceneVQA-G}.
Comments14 pages, 7 figures