arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1578 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1578 篇

2505.19769 2025-06-25 cs.RO cs.AI 57%

TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning

Yuhui Chen, Haoran Li, Zhennan Jiang, Haowei Wen, Dongbin Zhao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17705 2025-06-24 cs.CV 57%

DreamJourney: Perpetual View Generation with Video Diffusion Models

Bo Pan, Yang Chen, Yingwei Pan, Ting Yao, Wei Chen, Tao Mei

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Laboratory of Art and Archaeology Image(艺术与考古图像实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17608 2025-06-24 cs.CV 57%

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

Nikitha SR, Aradhya Neeraj Mathur, Tarun Ram Menta, Rishabh Jain, Mausoom Sarkar

机构 * Media and Data Science Research Lab, Adobe(Adobe媒体与数据科学研究实验室)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted in CVPR 2025 Workshop on What's Next in Multimodal Foundational Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17500 2025-06-24 cs.CV 57%

Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation

Julio Silva-Rodríguez, Fereshteh Shakeri, Houda Bahig, Jose Dolz, Ismail Ben Ayed

机构 * ÉTS Montréal(ÉTS蒙特利尔) Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM)(蒙特利尔大学中心医院研究中心(CRCHUM))

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments MICCAI 2025. Code: https://github.com/jusiro/SS-Text

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11748 2025-06-24 cs.CV 57%

ILIAS: Instance-Level Image retrieval At Scale

Giorgos Kordopatis-Zilos, Vladan Stojnić, Anna Manko, Pavel Šuma, Nikolaos-Antonios Ypsilantis, Nikos Efthymiadis, Zakaria Laskar, Jiří Matas, Ondřej Chum, Giorgos Tolias

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15903 2025-06-23 cs.LG 57%

VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics

Josef Kuchař, Marek Kadlčík, Michal Spiegel, Michal Štefánik

机构 * Kempelen Institute of Intelligent Technologies(智能技术研究所) Language Technology, University of Helsinki(语言技术,赫尔辛基大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05146 2025-06-23 cs.CV cs.CL 57%

CIVET: Systematic Evaluation of Understanding in VLMs

Massimo Rizzoli, Simone Alghisi, Olha Khomyn, Gabriel Roccabruna, Seyed Mahed Mousavi, Giuseppe Riccardi

机构 * Signals and Interactive Systems Lab, University of Trento, Italy(信号与交互系统实验室,特伦托大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15479 2025-06-19 cs.LG 57%

Creating User-steerable Projections with Interactive Semantic Mapping

Artur André Oliveira, Mateus Espadoto, Roberto Hirata, Roberto M. Cesar, Alex C. Telea

机构 * Institute of Mathematics and Statistics, University of São Paulo(数学与统计学研究所,圣保罗大学) Department of Information and Computing Sciences, Utrecht University(信息与计算科学系,乌得勒支大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13956 2025-06-19 cs.CV 57%

Improving LLM Video Understanding with 16 Frames Per Second

Yixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li, Zejun Ma, Chao Zhang

机构 * Tsinghua University(清华大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12826 2025-06-17 cs.CV 57%

LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling

Zhihan Zhang, Xiang Pan, Hongchen Wei, Zhenzhong Chen

机构 * School of Remote Sensing and Information Engineering, Wuhan University(武汉大学遥感与信息工程学院) School of Data Science, Lingnan University(岭南大学数据科学学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16458 2025-06-17 cs.RO cs.CV 57%

BiFold: Bimanual Cloth Folding with Language Guidance

Oriol Barbany, Adrià Colomé, Carme Torras

机构 * Institut de Robòtica i Informàtica Industrial, CSIC-UPC(工业机器人与计算机研究所,CSIC-UPC)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at ICRA 2025. Project page at https://barbany.github.io/bifold/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08817 2025-06-13 cs.CV 57%

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Shuyi Zhang, Xiaoshuai Hao, Yingbo Tang, Lingfeng Zhang, Pengwei Wang, Zhongyuan Wang, Hongxuan Ma, Shanghang Zhang

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artifcial Intelligence, UCAS(中国科学技术大学人工智能学院) Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院) Shenzhen International GraduateSchool,Tsinghua University(深圳国际研究生院,清华大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09175 2025-06-12 cs.CL cs.AI cs.SD eess.AS 57%

PHRASED: Phrase Dictionary Biasing for Speech Translation

Peidong Wang, Jian Xue, Rui Zhao, Junkun Chen, Aswin Shanmugam Subramanian, Jinyu Li

机构 * Microsoft USA(微软公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12559 2025-06-10 cs.CV cs.CL cs.MM 57%

AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu, Li Cao, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Huawei Technologies Co., Ltd.(华为技术有限公司) Shandong University(山东大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06173 2025-06-10 cs.CV 57%

VideoAuteur: Towards Long Narrative Video Generation

Junfei Xiao, Feng Cheng, Lu Qi, Liangke Gui, Jiepeng Cen, Zhibei Ma, Alan Yuille, Lu Jiang

机构 * Johns Hopkins University(约翰霍普金斯大学) ByteDance Project(字节跳动项目)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Preprint, https://videoauteur.github.io/; V2: Method is updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24870 2025-06-09 cs.CV 57%

GenSpace: Benchmarking Spatially-Aware Image Generation

Zehan Wang, Jiayang Xu, Ziang Zhang, Tianyu Pang, Chao Du, Hengshuang Zhao, Zhou Zhao

机构 * Zhejiang University(浙江大学) Sea AI Lab(海思人工智能实验室) The University of Hong Kong(香港大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05080 2025-06-06 cs.CL cs.CV 57%

Parking, Perception, and Retail: Street-Level Determinants of Community Vitality in Harbin

HaoTian Lan

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 22 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03511 2025-06-05 astro-ph.EP astro-ph.IM cs.AI eess.IV 57%

POLARIS: A High-contrast Polarimetric Imaging Benchmark Dataset for Exoplanetary Disk Representation Learning

Fangyi Cao, Bin Ren, Zihao Wang, Shiwei Fu, Youbin Mo, Xiaoyang Liu, Yuzhou Chen, Weixin Yao

机构 * UC Riverside(加州大学河滨分校) OCA/MPIA(天文台/马克斯·普朗克研究所) UT Chattanooga(田纳西大学查塔努加分校) MGH(麻省总医院) Harvard(哈佛大学) UC San Diego(加州大学圣地亚哥分校) Adobe(Adobe公司)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments 9 pages main text with 5 figures, 9 pages appendix with 9 figures. Submitted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15167 2025-06-05 cs.CV 57%

M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment

Chuan Cui, Kejiang Chen, Zhihua Wei, Wen Shen, Weiming Zhang, Nenghai Yu

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 24 pages. This work has been submitted to the ACM for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03414 2025-06-04 cs.CV cs.CL 57%

Enhancing Target-unspecific Tasks through a Features Matrix

Fangming Cui, Yonggang Zhang, Xuan Wang, Xinmei Tian, Jun Yu

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01902 2025-06-03 cs.CV cs.CL 57%

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination

Xinliu Zhong, Kayhan Batmanghelich, Li Sun

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 6 pages, 1 figure, accepted by 2024 IEEE Conference on Artificial Intelligence (CAI)

Journal ref 2024 IEEE Conference on Artificial Intelligence (CAI), 2024, 480-485

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01853 2025-06-03 cs.CV 57%

ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

Junliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie, Jun Zhu

机构 * Tsinghua University(清华大学) Peking University(北京大学) ShengShu(盛舒)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Project page: https://github.com/JAMESYJL/ShapeLLM-Omni

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13882 2025-06-03 cs.CV 57%

Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model

Long Le, Jason Xie, William Liang, Hung-Ju Wang, Yue Yang, Yecheng Jason Ma, Kyle Vedder, Arjun Krishna, Dinesh Jayaraman, Eric Eaton

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR 2025. Project website and open-source code: https://articulate-anything.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18672 2025-06-03 cs.CV 57%

FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models

Alice Heiman, Xiaoman Zhang, Emma Chen, Sung Eun Kim, Pranav Rajpurkar

机构 * Stanford University, USA(斯坦福大学) Harvard University, USA(哈佛大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24158 2025-06-02 cs.CV 57%

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Bo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song, Antoni B. Chan

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Baidu Inc.(百度公司) University of Sydney(悉尼大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22858 2025-05-30 cs.CV 57%

A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition

Sanjoy Kundu, Shanmukha Vellamcheti, Sathyanarayanan N. Aakur

机构 * CSSE Department, Auburn University(安全科学与工程系,阿伯丁大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Extended abstract of arXiv:2504.03948 for CVPR 2025 EgoVis Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22150 2025-05-30 cs.CV cs.CL 57%

Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging

Runze Xia, Shuo Feng, Renzhi Wang, Congchi Yin, Xuyun Wen, Piji Li

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(信息科技部模式分析与机器智能重点实验室) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CogSci2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22481 2025-05-29 stat.ML cs.LG 57%

Hypothesis Testing in Imaging Inverse Problems

Yiming Xi, Konstantinos Zygalakis, Marcelo Pereyra

机构 * School of Mathematical and Computer Sciences, Heriot-Watt University(赫瑞斯泰学院数学与计算机科学系,赫瑞瓦特大学) School of Mathematics, University of Edinburgh(爱丁堡大学数学学院) Maxwell Institute for Mathematical Sciences(麦克斯韦数学科学研究所)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13953 2025-05-29 cs.CL cs.AI 57%

Redundancy Principles for MLLMs Benchmarks

Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiaotong University(上海交通大学) Zhejiang University(浙江大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09441 2025-05-28 cs.CV eess.IV 57%

Structure-Accurate Medical Image Translation via Dynamic Frequency Balance and Knowledge Guidance

Jiahua Xu, Dawei Zhou, Lei Hu, Zaiyi Liu, Nannan Wang, Xinbo Gao

机构 * Xidian University(西安电子科技大学) Guangdong Provincial People’s Hospital(广东省人民医院) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

Comments Medical image translation, Diffusion model, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏