arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11402cs.CV

GroundSight团队参与GroundLM 2026共享任务:GoldenViewVQA

GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

发表机构山东大学软件学院 · 新加坡国立大学计算机学院
查看机构详情
  • School of Software, Shandong University(山东大学软件学院)
  • School of Computing, National University of Singapore(新加坡国立大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Kun Wang, Yupeng Hu, Ruping Cao, Hao Liu, Zhiran Li, Qianlong Xiang, Harry Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

GroundSight团队针对GoldenViewVQA任务提出无训练多阶段验证修正框架CoVeR-VQA,在官方测试集上取得84.75%联合准确率,经修正后达88.14%,凸显多视图推理中显式证据验证的重要性。

中文摘要 AI 辅助

GoldenViewVQA要求模型同时回答驾驶场景问题并定位包含支撑视觉证据的相机视角,精确的证据定位与答案正确性同等重要。本文提出CoVeR-VQA,这是一种用于 grounded 多视图VQA的无训练多阶段验证与修正框架。该框架从GPT-5.6的零样本预测出发,逐步应用基于Gemini-3.6-Flash的特定视角验证、基于Claude-Opus-5的先验引导联合验证,以及利用共享多视图场景的语义过滤问题组与验证集衍生先验知识的跨划分组级验证。在官方GoldenViewVQA测试集上,四阶段CoVeR-VQA流水线取得84.75%的联合准确率,较GPT-5.6零样本基线提升13.56个百分点,同时达到94.92%的答案准确率与86.44%的视角准确率。经两次额外的评估器指导的事后修正后,最终提交的运行结果取得88.14%的联合准确率。分析显示,支撑视角定位仍是剩余误差的主要来源,凸显显式证据验证对可靠多视图多模态推理的重要性。

英文摘要

GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for grounded multi-view VQA. Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge. On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75\% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92\% Answer Accuracy and 86.44\% View Accuracy. The final submitted run achieves 88.14\% Joint Accuracy after two additional evaluator-informed post-hoc corrections. Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.

↑