arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7464 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7464 篇

2511.13442 2025-11-19 cs.CV cs.AI 73%

Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline

Rui Zuo, Qinyue Tong, Zhe-Ming Lu, Ziqian Lu

机构 * Zhejiang University(浙江大学) Zhejiang Sci-Tech University(浙江科技学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00411 2025-11-18 cs.CV cs.AI 73%

Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis

Ran Tong, Jiaqi Liu, Tong Wang, Xin Hu, Su Liu, Lanruo Wang, Jiexi Xu

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) Independent Researcher(独立研究者) Duke University(杜克大学) University of Michigan Ann Arbor(密歇根大学安娜堡分校) Georgia Institute of Technology(佐治亚理工学院) University of California, Irvine(加州大学 Irvine 分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 6pages,3 figures.Uunder review of International Conference on Artificial Intelligence, Computer, Data Sciences and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03662 2025-11-17 cs.CV cs.AI cs.RO 73%

Zero-Shot Temporal Interaction Localization for Egocentric Videos

Erhang Zhang, Junyi Ma, Yin-Dong Zheng, Yixuan Zhou, Hesheng Wang

机构 * IRMV Lab, the Department of Automation, Shanghai Jiao Tong University(IRMV实验室,自动化系,上海交通大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09958 2025-11-14 cs.CV cs.AI 73%

Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification

Jeffrey Liu, Rongbin Hu

专题命中 视觉定位与Grounding :VLM(abstract);grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05565 2025-11-11 cs.CV cs.AI 73%

In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy

Shreyan Ganguly, Angona Biswas, Jaydeep Rade, Md Hasibul Hasan Hasib, Nabila Masud, Nitish Singla, Abhipsa Dash, Ushashi Bhattacharjee, Aditya Balu, Anwesha Sarkar, Adarsh Krishnamurthy, Soumik Sarkar

机构 * Iowa State University(爱荷华州立大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03757 2025-11-07 cs.LG cs.AI 73%

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

Xuan Ouyang, Senan Wang, Bouzhou Wang, Siyuan Xiahou, Jinrong Zhou, Yuekang Li

机构 * University of New South Wales(新南威尔士大学) University of Sydney(悉尼大学) The University of Hong Kong(香港大学) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27164 2025-11-03 cs.CV cs.AI 73%

Generating Accurate and Detailed Captions for High-Resolution Images

Hankyeol Lee, Gawon Seo, Kyounggyu Lee, Dogun Kim, Kyungwoo Song, Jiyoung Jung

机构 * Department of Artificial Intelligence, University of Seoul(首尔大学人工智能系) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系) Department of Applied Statistics, Yonsei University(延世大学应用统计系)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Work conducted in 2024; released for archival purposes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26151 2025-10-31 cs.CV cs.AI 73%

MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学部第一部门,LMU大学医院,慕尼黑,德国) Lunit Inc.(Lunit公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25616 2025-10-30 cs.LG cs.AI cs.RO 73%

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

Nikita Kachaev, Mikhail Kolosov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov

机构 * Cognitive AI Lab(认知人工智能实验室) Cognitive AI Lab, IAI MIPT(认知人工智能实验室,IAI MIPT)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13891 2025-10-17 cs.LG cs.AI 73%

K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding

Yifeng Yao, Yike Yun, Jing Wang, Huishuai Zhang, Dongyan Zhao, Ke Tian, Zhihao Wang, Minghui Qiu, Tao Wang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Bytedance(字节跳动)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17473 2025-10-17 cs.CV cs.AI cs.CL 73%

InfoDet: A Dataset for Infographic Element Detection

Jiangning Zhu, Yuxing Zhou, Zheng Wang, Juntao Yao, Yima Gu, Yuhui Yuan, Shixia Liu

机构 * BNRist, Tsinghua University(清华大学信息与技术研究院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10560 2025-10-14 cs.CL cs.AI cs.CV 73%

BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices

Euhid Aman, Esteban Carlin, Hsing-Kuo Pao, Giovanni Beltrame, Ghaluh Indah Permata Sari, Yie-Tarng Chen

机构 * NTUST(国立台湾科技大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments 6 pages, BabyLM Workshop, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10264 2025-10-14 cs.CV cs.AI 73%

MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs

Haonan Ge, Yiwei Wang, Ming-Hsuan Yang, Yujun Cai

机构 * Department of Computer Science and Engineering, University of California at Merced(计算机科学与工程系,加州大学默塞德分校)

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 73%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17374 2025-09-30 cs.CV cs.AI cs.IR 73%

From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge

Muhammad Tayyab Khan, Lequn Chen, Zane Yong, Jun Ming Tan, Wenhe Feng, Seung Ki Moon

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments Preprint submitted to Elsevier

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21356 2025-09-29 cs.CV cs.AI 73%

Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports

Razi Mahmood, Diego Machado-Reyes, Joy Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep Kalra, Ge Wang, Pingkun Yan, Tanveer Syeda-Mahmood

机构 * Rensselaer Polytechnic Institute, NY, USA(罗文学院) IBM Research, Almaden, CA, USA(IBM研究院) Stanford University, CA, USA(斯坦福大学) Massachusetts General Hospital (MGH), Boston, USA(麻省总医院)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments In proceedings MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19875 2025-09-25 cs.CV cs.AI 73%

Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13794 2025-09-22 cs.CV cs.AI 73%

LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation

Yang Zhou, Shiyu Zhao, Yuxiao Chen, Zhenting Wang, Can Jin, Dimitris N. Metaxas

机构 * Rutgers University(新泽西罗格斯大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13642 2025-09-18 cs.LG cs.CV 73%

LLM-I: LLMs are Naturally Interleaved Multimodal Creators

Zirun Guo, Feng Zhang, Kai Jia, Tao Jin

机构 * Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13234 2025-09-17 cs.AI cs.CV cs.HC 73%

Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy

Nadim Barakat, William Lotter

机构 * Dana-Farber Cancer Institute & Tufts University School of Medicine(达纳-法伯癌症研究所及塔夫茨大学医学院) Dana-Farber Cancer Institute Brigham and Women’s Hospital & Harvard Medical School(达纳-法伯癌症研究所布里特妇女医院及哈佛医学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 73%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 视觉定位与Grounding :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00284 2025-09-03 cs.CV cs.AI 73%

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

Liang Gong, Tommy, Wang, Sara Chaker, Yanchen Dong, Fouad Bousetouane, Brenden Morton, Mark Mendez

机构 * The University of Chicago(芝加哥大学) FabTrack

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12263 2025-09-01 cs.CV cs.AI 73%

Region-Level Context-Aware Multimodal Understanding

Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan, Debin Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18132 2025-08-26 cs.IR cs.AI cs.LG 73%

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang

机构 * Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以) NVIDIA(NVIDIA公司) Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09333 2025-08-12 cs.CV cs.AI 73%

Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Yufei Zhan, Shurong Zheng, Yousong Zhu, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 视觉定位与Grounding :vision language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025. Codes and datasets are released at https://github.com/jefferyZhan/Griffon

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16623 2025-07-23 cs.CV cs.LG 73%

Automatic Fine-grained Segmentation-assisted Report Generation

Frederic Jonske, Constantin Seibold, Osman Alperen Koras, Fin Bahnsen, Marie Bauer, Amin Dada, Hamza Kalisch, Anton Schily, Jens Kleesiek

机构 * Institute for AI in Medicine, University Medicine Essen(人工智能医学研究所,埃森大学医学中心)

专题命中 视觉定位与Grounding :LLaVA(abstract);grounding(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00151 2025-07-11 cs.CV cs.AI 73%

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Ser-Nam Lim, Rajiv Ramnath

机构 * Department of Computer Science(计算机科学系) Engineering, Ohio State University, Ohio, US(工程系,俄亥俄州立大学,俄亥俄,美国) Department of Computer Science, University of Central Florida, Florida, US(计算机科学系,中央佛罗里达大学,佛罗里达,美国)

专题命中 视觉定位与Grounding :visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21892 2025-06-30 cs.CV cs.AI 73%

SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation

Adam Goodge, Xun Xu, Bryan Hooi, Wee Siong Ng, Jingyi Liao, Yongyi Su, Xulei Yang

机构 * Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息通信研究所,科技研究局(A*STAR),新加坡) School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06184 2025-06-30 cs.CV cs.AI cs.CE cs.HC cs.MA 73%

PEACE: Empowering Geologic Map Holistic Understanding with MLLMs

Yangyu Huang, Tianyi Gao, Haoran Xu, Qihao Zhao, Yang Song, Zhipeng Gui, Tengchao Lv, Hao Chen, Lei Cui, Scarlett Li, Furu Wei

机构 * Microsoft Research(微软研究院) Chinese Academy of Geological Sciences(中国地质科学研究院) Wuhan University(武汉大学)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09174 2025-06-24 cs.CV cs.AI 73%

DART: An Automated End-to-End Object Detection Pipeline with Data Diversification, Open-Vocabulary Bounding Box Annotation, Pseudo-Label Review, and Model Training

Chen Xin, Andreas Hartel, Enkelejda Kasneci

专题命中 视觉定位与Grounding :InternVL(abstract);grounding(abstract);分类 cs.CV、cs.AI

Comments Corrected minor typos; no changes to results or conclusions

Journal ref Expert Systems with Applications 258 (2024): 125124

详情

展开后加载摘要…

URL PDF HTML 收藏