arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Inter-3D VQA:用于3D时空视觉问答的路侧多模态基准

Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering

Shaozu Ding, Linan Song, Dajiang Suo

arXiv 2608.28762首次发表:更新:

发表机构

The Polytechnic School, Arizona State University(亚利桑那州立大学理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出用于交叉口3D时空视觉问答的路侧多模态基准Inter-3D VQA,构建含40.7万问答对的数据集,提出基线模型Inter-Geo与评估框架Inter-Metrics,实验显示Inter-Geo在接地时空推理任务上优于基于图像的VLM。

AI 中文摘要

视觉问答(VQA)与多模态大语言模型(MLLM)的最新进展已支持针对交通场景的自然语言推理,但现有基准大多基于自车视角或2D路侧视频构建,限制了其评估对真实世界距离、轨迹、基础设施拓扑及安全关键交互的3D接地推理能力。本文提出Inter-3D VQA,这是一个面向交叉口的大规模路侧多模态3D时空视觉问答基准,由同步的点云和多视角图像构建,包含40.7万个问答对,覆盖车道级位置、物体关系、运动模式及近撞导向的交互推理。我们进一步提出Inter-Geo,一种整合物体级与场景级对齐的激光雷达表示的MLLM基线,以及Inter-Metrics,一种针对文本一致性、数值准确性和语义正确性的统一评估框架。实验表明,Inter-Geo优于基于图像的视觉语言模型(VLM),尤其在接地空间与时间推理任务上表现突出。我们的基准与代码可在https URL获取。

英文摘要

Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing benchmarks are largely built from ego-vehicle views or 2D roadside videos, limiting their ability to evaluate 3D-grounded reasoning over real-world distances, trajectories, infrastructure topology, and safety-critical interactions. We introduce Inter-3D VQA, a large-scale roadside multimodal benchmark for 3D spatiotemporally grounded VQA at intersections. Built from synchronized point clouds and multi-view images, Inter-3D VQA contains 407K QA pairs covering lane-level positions, object relationships, motion patterns, and near-miss-oriented interaction reasoning. We further propose Inter-Geo, an MLLM baseline that integrates object- and scene-level aligned LiDAR representations, and Inter-Metrics, a unified evaluation framework for textual consistency, numerical accuracy, and semantic correctness. Experiments show that Inter-Geo outperforms image-based VLMs, especially on grounded spatial and temporal reasoning tasks. Our benchmark and codes are available at https://github.com/ASU-Suo-Lab/Inter-3D-VQA .

CommentsAccepted to EMNLP 2026 main conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑