arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2408.15172 2025-10-24 cs.IR cs.CL cs.CV 57%

X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation

Hanjia Lyu, Ryan Rossi, Xiang Chen, Md Mehrab Tanjim, Stefano Petrangeli, Somdeb Sarkhel, Jiebo Luo

机构 * University of Rochester(罗切斯特大学) Adobe Research(Adobe研究)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13061 2025-10-23 cs.CV 57%

3D Visual Illusion Depth Estimation

Chengtang Yao, Zhidan Liu, Jiaxi Zeng, Lidong Yu, Yuwei Wu, Yunde Jia

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学,中国) Guangdong Provincial Key Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, Shenzhen, China(广东省机器感知与智能计算重点实验室,深圳MSU-BIT大学,深圳,中国) NVIDIA

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

Comments NeurIPS 2025, Project: https://github.com/YaoChengTang/3D-Visual-Illusion-Depth-Estimation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18837 2025-10-22 cs.CV 57%

FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning

Yubin Zheng, Pak-Hei Yeung, Jing Xia, Tianjie Ju, Peng Tang, Weidong Qiu, Jagath C. Rajapakse

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18726 2025-10-22 cs.CV 57%

IF-VidCap: Can Video Caption Models Follow Instructions?

Shihao Li, Yuanxing Zhang, Jiangtao Wu, Zhide Lei, Yiwen He, Runzhe Wen, Chenxi Liao, Chengkang Jiang, An Ping, Shuo Gao, Suhan Wang, Zhaozhou Bian, Zijun Zhou, Jingyi Xie, Jiayi Zhou, Jing Wang, Yifan Yao, Weihao Xie, Yingshui Tan, Yanghai Wang, Qianqian Xie, Zhaoxiang Zhang, Jiaheng Liu

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments https://github.com/NJU-LINK/IF-VidCap

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04641 2025-10-22 cs.LG math.ST stat.ML stat.TH 57%

A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI

Kazusato Oko, Licong Lin, Yuhang Cai, Song Mei

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16034 2025-10-21 cs.CV 57%

VisualLens: Personalization through Task-Agnostic Visual History

Wang Bill Zhu, Deqing Fu, Kai Sun, Yi Lu, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Yue Liu, Anuj Kumar, Xin Luna Dong

机构 * Meta University of Southern California(南加州大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12444 2025-10-15 cs.CV 57%

A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation

Shaoyang Zhou, Yingshu Li, Yunyi Liu, Lingqiao Liu, Lei Wang, Luping Zhou

机构 * School of Electrical and Computer Engineering, The University of Sydney(电气与计算机工程学院,悉尼大学) School of Computer Science, The University of Adelaide(计算机科学学院,阿德莱德大学) School of Computing and Information Technology, University of Wollongong(计算与信息科技学院,沃伦冈大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12119 2025-10-15 cs.CV 57%

ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation

Ziyuan Luo, Yangyi Zhao, Ka Chun Cheung, Simon See, Renjie Wan

机构 * Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) NVIDIA AI Technology Center, NVIDIA(NVIDIA 人工智能技术中心)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11498 2025-10-14 cs.LG cs.CL 57%

ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding

Yuhang Li, Chenchen Zhang, Ruilin Lv, Ao Liu, Ken Deng, Yuanxing Zhang, Jiaheng Liu, Wiggin Zhou, Bo Zhou

机构 * LLM Department, Tencent(腾讯大语言模型部门) Peking University(北京大学) Nanjing University(南京大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07501 2025-10-14 cs.CV cs.HC 57%

FormCoach: Lift Smarter, Not Harder

Xiaoye Zuo, Nikos Athanasiou, Ginger Delmas, Yiming Huang, Xingyu Fu, Lingjie Liu

机构 * University of Pennsylvania(宾夕法尼亚大学) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03948 2025-10-13 cs.CV 57%

ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition

Sanjoy Kundu, Shanmukha Vellamcheti, Sathyanarayanan N. Aakur

机构 * CSSE Department, Auburn University(计算机科学与工程系,阿伯丁大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. 17 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09078 2025-10-13 cs.GR cs.LG 57%

MCMC: Bridging Rendering, Optimization and Generative AI

Gurprit Singh, Wenzel Jakob

机构 * Max Planck Institute for Informatics(马克斯·普朗克研究所信息学研究所) EPFL(瑞士联邦理工学院)

专题命中 其他VLM :vision language model(abstract);分类 cs.LG

Comments SIGGRAPH Asia 2024 Courses. arXiv admin note: text overlap with arXiv:2208.11970 by other authors

Journal ref SIGGRAPH Asia 2024 Courses, Article No.: 8, Pages 1 - 27

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17132 2025-10-08 cs.AI 57%

Applications of Large Models in Medicine

YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu, Yuxin Jia, Lujia Ge, Jing-min Yang

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04257 2025-10-07 cs.CR cs.AI 57%

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao

机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02815 2025-10-06 cs.CV 57%

Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis

Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所) Chinese Academy of Sciences(中国科学院) The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR2026 under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02787 2025-10-06 cs.CV 57%

OTR: Synthesizing Overlay Text Dataset for Text Removal

Jan Zdenek, Wataru Shimoda, Kota Yamaguchi

机构 * CyberAgent Tokyo Japan(CyberAgent东京日本)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01247 2025-10-03 cs.CL cs.AI 57%

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

Punit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, Jose G Moreno

机构 * Indian Institute of Technology Patna(印度理工学院帕纳布分校) Sardar Patel Institute of Technology(萨达尔·帕特尔技术学院) Universitas Gadjah Mada(加查马大学) King Mongkut’s Institute of Technology Ladkrabang(拉差班国王技术学院) Shenzhen Technology University(深圳技术大学) Université de Toulouse(图卢兹大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments 52 pages, 56 figures; appearing at EMNLP'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01185 2025-10-02 cs.LG 57%

Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs

Leyla Mirvakhabova, Babak Ehteshami Bejnordi, Gaurav Kumar, Hanxue Liang, Wanru Zhao, Paul Whatmough

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25297 2025-10-02 cs.SE cs.AI 57%

Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development

Yuxuan Wan, Tingshuo Liang, Jiakai Xu, Jingyu Xiao, Yintong Huo, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学) Columbia University in the City of New York(哥伦比亚大学) Singapore Management University(新加坡管理学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00806 2025-10-02 cs.CV 57%

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心、自动化研究所、中国科学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国) School of Artificial Intelligence, University of Chinese Academy of Science, Beijing, China(人工智能学院、中国科学院大学、北京中国) Wuhan AI Research, Wuhan, China(武汉人工智能研究、武汉中国) MAPLE Lab, Westlake University(MAPLE实验室、西湖大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26555 2025-10-01 cs.CV 57%

Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation

Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani

机构 * Stability AI Arizona State University(亚利桑那州立大学) Google DeepMind(谷歌DeepMind)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments NeurIPS 2025. Project Page : https://stable-cinemetrics.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25863 2025-10-01 cs.CV 57%

MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image Classification

Junjie Zhou, Wei Shao, Yagao Yue, Wei Mu, Peng Wan, Qi Zhu, Daoqiang Zhang

机构 * The College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25817 2025-10-01 cs.CL cs.CV 57%

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang

机构 * Teamreboott Inc.(Teamreboott公司) MIRI D.I.H Inc.(MIRI D.I.H公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23242 2025-09-30 cs.CV 57%

TATTOO: Training-free AesTheTic-aware Outfit recOmmendation

Yuntian Wu, Xiaonan Hu, Ziqi Zhou, Hao Lu

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21980 2025-09-29 cs.CV 57%

Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm

Zeyu Wang, Baiyu Chen, Kun Yan, Hongjing Piao, Hao Xue, Flora D. Salim, Yuanchun Shi, Yuntao Wang

机构 * Key Laboratory of Pervasive Computing, Tsinghua University(清华大学普适计算重点实验室) The University of New South Wales(新南威尔士大学) SKLSDE Lab, Beihang University(北航SKLSDE实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15472 2025-09-29 cs.CV 57%

Efficient Multimodal Dataset Distillation via Generative Models

Zhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Central Florida(中央佛罗里达大学) Cisco Research(思科研究)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16663 2025-09-26 cs.CL cs.AI 57%

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han, Zhiheng Wang, Xinyu Liu, Yingya Zhang, Shiwei Zhang, Difan Zou

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

Comments 31 pages, 16 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20279 2025-09-25 cs.CV q-bio.QM 57%

A co-evolving agentic AI system for medical imaging analysis

Songhao Li, Jonathan Xu, Tiancheng Bao, Yuxuan Liu, Yuchen Liu, Yihang Liu, Lilin Wang, Wenhui Lei, Sheng Wang, Yinuo Xu, Yan Cui, Jialu Yao, Shunsuke Koga, Zhi Huang

机构 * Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学) Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学) The Wharton School, University of Pennsylvania(沃顿商学院,宾夕法尼亚大学) Department of Bioengineering, University of Pennsylvania(生物工程系,宾夕法尼亚大学) Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学) Department of Biostatistics, Epidemiology & Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16560 2025-09-23 cs.CV 57%

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

Ji Soo Lee, Byungoh Ko, Jaewon Cho, Howoong Lee, Jaewoon Byun, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Hanwha Vision(翰威英航) KAIST(韩国科学技术院)

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13939 2025-09-18 cs.CV 57%

Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation

Gia Khanh Nguyen, Yifeng Huang, Minh Hoai

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Stony Brook University(石溪大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏