arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-12 至 2025-09-12 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 4 篇

2505.19455 2025-09-12 cs.CV cs.AI cs.LG 82%

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

Xu Li, Fan Lyu

机构 * Khoury College of Computer Sciences, Northeastern University(东北大学克劳尔计算机科学学院) New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07084 2025-09-12 cs.RO 82%

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

Shucheng Huang, Freda Shi, Chen Sun, Jiaming Zhong, Minghao Ning, Yufeng Yang, Yukun Lu, Hong Wang, Amir Khajepour

机构 * MVSLab, Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系MVSLab) CompLING Lab, David R. Cheriton School of Computer Science, University of Waterloo(滑铁卢大学大卫·R·切里顿计算机科学学院CompLING Lab) Department of Data and Systems Engineering, University of Hong Kong(香港大学数据与系统工程系) Department of Mechanical Engineering, University of New Brunswick(新不伦瑞克大学机械工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)

专题命中 视觉问答 :multimodal large language model(title,abstract);MLLM(abstract)

Comments This work has been accepted to IEEE Transactions on Vehicular Technology. Please refer to the copyright notice for additional information

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09159 2025-09-12 cs.CV cs.AI 81%

A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering

Zhiyue Liu, Sihang Liu, Jinyuan Liu, Xinru Zhang

机构 * School of Computer, Electronics and Information(计算机、电子与信息学院) Guangxi University(广西大学) Guangxi Key Laboratory of Multimedia Communications and Network Technology(广西多媒体通信与网络技术重点实验室)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by the IEEE International Conference on Multimedia and Expo (ICME 2025) for oral presentation. © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09254 2025-09-12 cs.CV cs.MM 70%

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) National University of Singapore(新加坡国立大学) CVTE Sun Yat-sen University(孙中山大学) Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科) Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科) School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments 40 pages, 26 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏