Vision-Based Natural Language Scene Understanding for Autonomous Driving: An Extended Dataset and a New Model for Traffic Scene Description Generation
基于视觉的自然语言场景理解用于自动驾驶:一个扩展数据集和一个用于交通场景描述生成的新模型
机构 * Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系) ; School of Electrical and Computer Engineering, College of Engineering, University of Tehran(德黑兰大学电气与计算机工程学院)
专题命中 仿真评测 :autonomous driving(title);分类 cs.CV、cs.AI
AI总结 本文提出了一种基于视觉的自然语言场景理解模型,通过扩展数据集和混合注意力机制,提升自动驾驶中交通场景描述生成的准确性和丰富性。
Comments Under review at Computer Vision and Image Understanding (submitted July 25, 2025)