发表机构
The Polytechnic School, Arizona State University(亚利桑那州立大学理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出用于交叉口3D时空视觉问答的路侧多模态基准Inter-3D VQA,构建含40.7万问答对的数据集,提出基线模型Inter-Geo与评估框架Inter-Metrics,实验显示Inter-Geo在接地时空推理任务上优于基于图像的VLM。
AI 中文摘要
视觉问答(VQA)与多模态大语言模型(MLLM)的最新进展已支持针对交通场景的自然语言推理,但现有基准大多基于自车视角或2D路侧视频构建,限制了其评估对真实世界距离、轨迹、基础设施拓扑及安全关键交互的3D接地推理能力。本文提出Inter-3D VQA,这是一个面向交叉口的大规模路侧多模态3D时空视觉问答基准,由同步的点云和多视角图像构建,包含40.7万个问答对,覆盖车道级位置、物体关系、运动模式及近撞导向的交互推理。我们进一步提出Inter-Geo,一种整合物体级与场景级对齐的激光雷达表示的MLLM基线,以及Inter-Metrics,一种针对文本一致性、数值准确性和语义正确性的统一评估框架。实验表明,Inter-Geo优于基于图像的视觉语言模型(VLM),尤其在接地空间与时间推理任务上表现突出。我们的基准与代码可在https URL获取。
英文摘要
Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing benchmarks are largely built from ego-vehicle views or 2D roadside videos, limiting their ability to evaluate 3D-grounded reasoning over real-world distances, trajectories, infrastructure topology, and safety-critical interactions. We introduce Inter-3D VQA, a large-scale roadside multimodal benchmark for 3D spatiotemporally grounded VQA at intersections. Built from synchronized point clouds and multi-view images, Inter-3D VQA contains 407K QA pairs covering lane-level positions, object relationships, motion patterns, and near-miss-oriented interaction reasoning. We further propose Inter-Geo, an MLLM baseline that integrates object- and scene-level aligned LiDAR representations, and Inter-Metrics, a unified evaluation framework for textual consistency, numerical accuracy, and semantic correctness. Experiments show that Inter-Geo outperforms image-based VLMs, especially on grounded spatial and temporal reasoning tasks. Our benchmark and codes are available at https://github.com/ASU-Suo-Lab/Inter-3D-VQA .
CommentsAccepted to EMNLP 2026 main conference