RAU: Reference-based Anatomical Understanding with Vision Language Models
Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun
机构
*
United Imaging Intelligence(联合影像智能)
;
School of Computing, University of Georgia(佐治亚大学计算机学院)
;
Department of Electrical and Computer Engineering, Northwestern University(西北大学电气与计算机工程系)
专题命中
视觉问答
:vision language model(title);vision-language model(abstract);VLM(abstract);visual reasoning(abstract)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
Maximiliano Hormazábal Lagos, Héctor Cerezo-Costas, Dimosthenis Karatzas
机构
*
Computer Vision Center, Universitat Autònoma de Barcelona(计算机视觉中心,巴塞罗那自治大学)
;
Gradiant
专题命中
视觉问答
:vision language model(title,abstract);visual question answering(title);分类 cs.CV、cs.AI、cs.LG
CommentsThis work has been accepted for presentation at the 16th Conference and Labs of the Evaluation Forum (CLEF 2025) and will be published in the proceedings by Springer in the Lecture Notes in Computer Science (LNCS) series. Please cite the published version when available
机构
*
Kyoto University(京都大学)
;
NII LLMC(国立信息学研究所LLMC)
;
Waseda University(早稻田大学)
;
Institute of Science Tokyo(东京科学大学)
;
NII(国立信息学研究所)
;
Aichi Institute of Technology(爱知工业大学)
;
Institute of Physical and Chemical Research(理化学研究所)
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery
DM-KG:一种提升街景图像中视觉语言模型空间认知的新方法
Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, Zixuan Li, Yi Cheng
机构
*
Institute of Surveying and Mapping, Information Engineering University(信息工程大学测绘学院)
;
Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences(中国科学院地理科学与资源研究所)
Search-based Testing of Vision Language Models for In-Car Scene Understanding
基于搜索的车内场景理解视觉语言模型测试
Lev Sorokin, Chen Yang, Ken E. Friedl, Andrea Stocco
机构
*
BMW Group, Technical University of Munich(宝马集团、慕尼黑技术大学)
;
Technical University of Munich(慕尼黑技术大学)
;
Technical University of Munich, fortiss GmbH(慕尼黑技术大学、fortiss GmbH)
专题命中
视觉问答
:VLM(summary_cn,abstract_cn);vision language model(title);vision-language model(abstract);分类 cs.CV
机构
*
Arizona State University(亚利桑那州立大学)
;
Clemson University(克莱姆森大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
University of Notre Dame(诺特丹大学)
;
Florida State University(佛罗里达州立大学)
;
Rice University(里德大学)
;
NVIDIA(英伟达)
;
Mayo Clinic(梅奥诊所)
专题命中
视觉问答
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);visual question answering(abstract);分类 cs.CV
机构
*
Mehta Family School of Data Science and Artificial Intelligence(梅hta家族数据科学与人工智能学院)
;
Indian Institute of Technology, Roorkee(印度理工学院罗奥克学院)
;
Department of Civil Engineering(土木工程系)
;
Department of Electronics and Communication(电子与通信系)
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving
NuRisk:面向自动驾驶中 agent 级风险评估的视觉问答数据集
Yuan Gao, Mattia Piccinini, Roberto Brusnicki, Yuchen Zhang, Johannes Betz
机构
*
Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,TUM工程与设计学院,慕尼黑技术大学)
;
Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所(MIRMI))
专题命中
视觉问答
:visual question answering(title,abstract);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.AI
机构
*
School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学)
;
School of Future Technology, South China University of Technology(未来技术学院,华南理工大学)
;
Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)