arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2408.08396 2025-05-27 cs.CV cs.CL 57%

Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment

Daniele Rege Cambrin, Gabriele Scaffidi Militone, Luca Colomba, Giovanni Malnati, Daniele Apiletti, Paolo Garza

机构 * Politecnico di Torino(托里尼理工大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at ECCV 2024 CV2 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01263 2025-05-27 cs.CV cs.CL 57%

Generalizable Prompt Learning of CLIP: A Brief Overview

Fangming Cui, Yonggang Zhang, Xuan Wang, Xule Wang, Liang Xiao

机构 * Meituan(美团)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12746 2025-05-26 cs.AI 57%

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs

Haruka Asanuma, Naoko Koide-Majima, Ken Nakamura, Takato Horii, Shinji Nishimoto, Masafumi Oizumi

机构 * The University of Tokyo, Graduate School of Arts and Sciences(东京大学艺术与科学研究生院) Center for Information and Neural Networks (CiNet), National Institute of Information and Communications Technology(信息与神经网络中心(CiNet),信息与通信技术国家研究所) The University of Osaka, Graduate School of Frontier Biosciences(大阪大学前沿生命科学研究生院) The University of Tokyo, Faculty of Engineering(东京大学工学部) The University of Osaka, Graduate School of Engineering Science(大阪大学工学研究院) The University of Osaka, Graduate School of Medicine(大阪大学医学研究院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15415 2025-05-23 cs.CV 57%

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery

Siqi Fan, Yuguang Xie, Bowen Cai, Ailin Xie, Gaochao Liu, Mu Qiao, Jie Xing, Zaiqing Nie

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究所(AIR),清华大学) PharMolix Inc.(PharMolix公司)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15145 2025-05-22 cs.CV 57%

CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

Xinran Wang, Songyu Xu, Xiangxuan Shan, Yuxuan Zhang, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) China Mobile Research Institute(中国移动研究院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13669 2025-05-21 cs.CV cs.RO 57%

GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching

Barkin Dagda, Muhammad Awais, Saber Fallah

机构 * Connected and Autonomous Vehicles Lab (CAV-Lab)(连接与自动驾驶车辆实验室) University of Surrey(萨里大学) Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音和信号处理中心)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13082 2025-05-20 cs.SD cs.AI eess.AS 57%

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

Kyeongman Park, Seongho Joo, Kyomin Jung

机构 * Seoul National University(首尔国立大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12660 2025-05-20 cs.CV 57%

Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps

Ziqi Wen, Jonathan Skaza, Shravan Murlidaran, William Y. Wang, Miguel P. Eckstein

机构 * Department of Computer Science University of California Santa Barbara(计算机科学系加州大学圣芭芭拉分校) Graduate Program in Dynamical Neuroscience University of California Santa Barbara(动态神经科学研究生项目加州大学圣芭芭拉分校) Department of Psychological and Brain Sciences University of California Santa Barbara(心理学与脑科学系加州大学圣芭芭拉分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13372 2025-05-20 cs.GR cs.CV 57%

MoVer: Motion Verification for Motion Graphics Animations

Jiaju Ma, Maneesh Agrawala

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to ACM Transactions on Graphics (SIGGRAPH 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11865 2025-05-19 cs.CV 57%

From Image to Video, what do we need in multimodal LLMs?

Suyuan Huang, Haoxin Zhang, Linqing Zhong, Honggu Chen, Yan Gao, Yao Hu, Zengchang Qin

机构 * Intelligent Computing and Machine Learning Lab, School of ASEE, Beihang University(北京航空航天大学自动化学院智能计算与机器学习实验室) Xiaohongshu(小红书) School of Sino-French Engineer, Beihang University(北京航空航天大学中法工程师学院) College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09659 2025-05-16 cs.LG cs.CL 57%

LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language Models

Long Chen, Xiaotian Song, Yanan Sun

机构 * Collage of Computer Science, Sichuan University(计算机科学学院,四川大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07347 2025-05-13 cs.CV 57%

AI-Enabled Accurate Non-Invasive Assessment of Pulmonary Hypertension Progression via Multi-Modal Echocardiography

Jiewen Yang, Taoran Huang, Shangwei Ding, Xiaowei Xu, Qinhua Zhao, Yong Jiang, Jiarong Guo, Bin Pu, Jiexuan Zheng, Caojin Zhang, Hongwen Fei, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) Guangdong Cardiovascular Institute, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University(广东省心血管病研究所,广东省人民医院(广东省医学科学院)) Department of Ultrasound, The First Affiliated Hospital of Guangzhou Medical University(广州市第一人民医院超声科) Department of Pulmonary Circulation, Shanghai Pulmonary Hospital, Tongji University School of Medicine(上海 pulmonary 医院,同济大学医学院) Department of Echocardiography, Fuwai Hospital Chinese Academy of Medical Sciences(阜外医院中国医学科学院) Guangdong Provincial Key Laboratory of South China Structural Heart Disease(广东省南方结构性心脏病重点实验室) State Key Laboratory of Cardiovascular Disease, Department of Echocardiography, National Center for Cardiovascular Diseases, Fuwai Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(心血管疾病国家重点实验室,国家心血管病中心,阜外医院,中国医学科学院和北京协和医学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03611 2025-05-07 cs.CV 57%

Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images

Fangling Jiang, Qi Li, Weining Wang, Wei Shen, Bing Liu, Zhenan Sun

机构 * School of Computer Science, University of South China(南方科技大学计算机科学学院) New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA) School of Artificial Intelligence, UCAS(人工智能学院,UCAS) The Laboratory of Cognition and Decision Intelligence for Complex Systems, CASIA(复杂系统认知与决策智能实验室,CASIA) OPPO AI Center(OPPO AI中心)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00569 2025-05-02 cs.CV 57%

AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis

Enmin Zhong, Carlos R. del-Blanco, Daniel Berjón, Fernando Jaureguizar, Narciso García

机构 * Grupo de Tratamiento de Imágenes (GTI), Information Processing and Telecommunications Center, ETSI Telecomunicación, Universidad Politécnica de Madrid(图像处理小组(GTI)、信息处理与电信中心、电信工程学院、马德里理工大学)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

Comments 6 pages, 3 figures,Accepted for the poster session at the CV4Animals workshop: Computer Vision for Animal Behavior Tracking and Modeling In conjunction with Computer Vision and Pattern Recognition 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07072 2025-04-30 cs.CL cs.CV 57%

Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar, Shivalika Singh, Fabian Farestam, Angelika Romanou, Danylo Boiko, Dipika Khullar, Mike Zhang, Dominik Krzemiński, Jekaterina Novikova, Luísa Shimabucoro, Joseph Marvin Imperial, Rishabh Maheshwary, Sharad Duwal, Alfonso Amayuelas, Swati Rajwal, Jebish Purbey, Ahmed Ruby, Nicholas Popovič, Marek Suppa, Azmine Toushik Wasi, Ram Mohan Rao Kadiyala, Olga Tsymboi, Maksim Kostritsya, Bardia Soltani Moakhar, Gabriel da Costa Merlin, Otávio Ferracioli Coletti, Maral Jabbari Shiviari, MohammadAmin farahani fard, Silvia Fernandez, María Grandury, Dmitry Abulkhanov, Drishti Sharma, Andre Guarnier De Mitri, Leticia Bossatto Marchezi, Setayesh Heydari, Johan Obando-Ceron, Nazar Kohut, Beyza Ermis, Desmond Elliott, Enzo Ferrante, Sara Hooker, Marzieh Fadaee

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments v2: corrected the author list

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19742 2025-04-29 cs.CV 57%

EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia

Valerie Zermatten, Javiera Castillo-Navarro, Pallavi Jain, Devis Tuia, Diego Marcos

机构 * EPFL(瑞士联邦理工学院) CNAM(法国国家科学与技术研究中心) INRIA(法国国家信息与自动化研究所) CIHEAM-IAMM(CIHEAM- IAMM) Univ. of Montpellier(蒙彼利埃大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

Comments Accepted at EarthVision 2025 (CVPRW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18961 2025-04-29 cs.IR cs.AI 57%

Feature Fusion Revisited: Multimodal CTR Prediction for MMCTR Challenge

Junjie Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments A technical report for the MMCTR Challenge held by EReL@MIR Workshop at WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18509 2025-04-28 cs.CV 57%

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

Shivam Duggal, Yushi Hu, Oscar Michel, Aniruddha Kembhavi, William T. Freeman, Noah A. Smith, Ranjay Krishna, Antonio Torralba, Ali Farhadi, Wei-Chiu Ma

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments CVPR 2025. Project page and codes: https://eval3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16433 2025-04-24 cs.CV 57%

FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing

Hariseetharam Gunduboina, Muhammad Haris Khan, Biplab Banerjee

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔分校) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13971 2025-04-22 cs.CY cs.AI cs.ET cs.NI 57%

The Future of Internet of Things and Multimodal Language Models in 6G Networks: Opportunities and Challenges

Abdelrahman Soliman

机构 * University of Guelph(圭尔夫大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13650 2025-04-21 cs.CV 57%

EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model

Sijing Li, Tianwei Lin, Lingshuai Lin, Wenqiao Zhang, Jiang Liu, Xiaoda Yang, Juncheng Li, Yucheng He, Xiaohui Song, Jun Xiao, Yueting Zhuang, Beng Chin Ooi

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13209 2025-04-21 cs.CR cs.AI 57%

On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks

Ting Bi, Chenghang Ye, Zheyu Yang, Ziyi Zhou, Cui Tang, Jun Zhang, Zui Tao, Kailong Wang, Liting Zhou, Yang Yang, Tianlong Yu

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11509 2025-04-18 cs.IR cs.CV 57%

PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage

Wenyi Zhang, Ju Jia, Xiaojun Jia, Yihao Huang, Xinfeng Li, Cong Wu, Lina Wang

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10471 2025-04-15 cs.CV cs.CL 57%

MIEB: Massive Image Embedding Benchmark

Chenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling, Xin Zhang, Márton Kardos, Roman Solomatin, Noura Al Moubayed, Kenneth Enevoldsen, Niklas Muennighoff

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09598 2025-04-15 cs.CV 57%

DualPrompt-MedCap: A Dual-Prompt Enhanced Approach for Medical Image Captioning

Yining Zhao, Ali Braytee, Mukesh Prasad

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 11 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14911 2025-04-15 cs.CV 57%

Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

Siyuan Yan, Ming Hu, Yiwen Jiang, Xieji Li, Hao Fei, Philipp Tschandl, Harald Kittler, Zongyuan Ge

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Our dataset and code will be publicly available at https://github.com/SiyuanYan1/Derm1M

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13878 2025-04-15 cs.HC cs.CV 57%

Eye Gaze as a Signal for Conveying User Attention in Contextual AI Systems

Ethan Wilson, Naveen Sendhilnathan, Charlie S. Burlingham, Yusuf Mansour, Robert Cavin, Sai Deep Tetali, Ajoy Savio Fernandes, Michael J. Proulx

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

Comments To appear in ETRA '25: Proceedings of the 2025 Symposium on Eye Tracking Research and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00672 2025-04-15 cs.CV 57%

ExpertAF: Expert Actionable Feedback from Video

Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, Kristen Grauman

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07954 2025-04-11 cs.CV cs.CL 57%

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

En Yu, Kangheng Lin, Liang Zhao, Jisheng Yin, Yana Wei, Yuang Peng, Haoran Wei, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang, Wenbing Tao

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments Github page: https://github.com/linkangheng/PR1

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07643 2025-04-11 cs.IR cs.CL cs.CV 57%

CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections

Florian Schneider, Narges Baba Ahmadi, Niloufar Baba Ahmadi, Iris Vogel, Martin Semmann, Chris Biemann

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏