arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2510.20696 2025-10-24 cs.CV 57%

Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward

Jing Bi, Guangyu Sun, Ali Vosoughi, Chen Chen, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21401 2025-10-24 cs.CV 57%

JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation

Md Jueal Mia, M. Hadi Amini

机构 * Knight Foundation School of Computing and Information Sciences (KFSCIS)(骑士基金会计算与信息科学学院) Florida International University(佛罗里达国际大学) Sustainability, Optimization, and Learning for InterDependent networks laboratory (solid lab)(可持续性、优化与互依赖网络学习实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19802 2025-10-23 cs.CV 57%

Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models

Xiaozhen Qiao, Jingkai Zhao, Yuqiu Jiang, Xianda Guo, Zhe Sun, Hongyuan Zhang, Xuelong Li

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Institute of Artificial Intelligence (TeleAI), China Telecom, P. R. China(人工智能研究所(TeleAI),中国电信,中华人民共和国) College of Computer Science, Wuhan University(计算机科学学院,武汉大学) School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院,光学与电子学(iOPEN),西北工业大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16814 2025-10-23 cs.LG cs.CV 57%

Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning

Junhao Shen, Haiteng Zhao, Yuzhe Gu, Songyang Gao, Kuikun Liu, Haian Huang, Jianfei Gao, Dahua Lin, Wenwei Zhang, Kai Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18262 2025-10-22 cs.CV 57%

UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding

Da Zhang, Chenggang Rong, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI)、中国电信)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments We have released V1, which only reports the test results. Our work is still ongoing, and the next version will be coming soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17777 2025-10-21 cs.CV 57%

SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference

Samir Khaki, Junxian Guo, Jiaming Tang, Shang Yang, Yukang Chen, Konstantinos N. Plataniotis, Yao Lu, Song Han, Zhijian Liu

机构 * NVIDIA MIT(麻省理工学院) UC San Diego(南加州大学圣地亚哥分校) University of Toronto(多伦多大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09607 2025-10-20 cs.CV 57%

VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation

Shaoqi Dong, Chaoyou Fu, Haihan Gao, Yi-Fan Zhang, Chi Yan, Chu Wu, Xiaoyu Liu, Yunhang Shen, Jing Huo, Deqiang Jiang, Haoyu Cao, Yang Gao, Xing Sun, Ran He, Caifeng Shan

机构 * Nanjing University(南京大学) Tencent Youtu Lab(腾讯优图实验室) CASIA(中国科学院自动化研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Homepage: https://ltbai.github.io/VITA-VLA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01842 2025-10-20 cs.LG cs.AI 57%

GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping

Qifu Wen, Xi Zeng, Zihan Zhou, Shuaijun Liu, Mehdi Hosseinzadeh, Ningxin Su, Reza Rawassizadeh

机构 * Department of Computer Science, Boston University Metropolitan College(波士顿大学计算机科学系) Information Hub, The Hong Kong University of Science and Technology, Guangzhou(香港科技大学广州信息中心) School of Engineering and Technology, Duy Tan University, Da Nang, Vietnam(杜益大学工程科技学院,岘港,越南) Department of AI, School of Computer Science and Engineering, Galgotias University, Greater Noida, India(加洛吉亚大学人工智能系,诺伊达,印度)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12126 2025-10-17 cs.CV 57%

MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites

Zhenxin Lei, Zhangwei Gao, Changyao Tian, Erfei Cui, Guanzhou Chen, Danni Yang, Yuchen Duan, Zhaokai Wang, Wenhao Li, Weiyun Wang, Xiangyu Zhao, Jiayi Ji, Yu Qiao, Wenhai Wang, Gen Luo

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) University of Chinese Academy of Science(中国科学院大学) Xiamen University(厦门大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16411 2025-10-17 cs.CV cs.LG 57%

Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression

Sreetama Sarkar, Yue Che, Alex Gavin, Peter A. Beerel, Souvik Kundu

机构 * University of Southern California(美国南加州大学) Intel Labs(英特尔实验室) Harvard-Westlake School(哈佛-韦斯利学校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13359 2025-10-16 cs.IR cs.CV cs.LG 57%

Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models

Yuki Yada, Sho Akiyama, Ryo Watanabe, Yuta Ueno, Yusuke Shido, Andre Rusli

机构 * Mercari, Inc.(Mercari公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ACM RecSys 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13190 2025-10-16 cs.CL 57%

SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10962 2025-10-14 cs.LG cs.AI 57%

MC#: Mixture Compressor for Mixture-of-Experts Large Models

Wei Huang, Yue Liao, Yukang Chen, Jianhui Liu, Haoru Tan, Si Liu, Shiming Zhang, Shuicheng Yan, Xiaojuan Qi

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系) School of Computing, National University of Singapore(新加坡国立大学计算机学院) NVIDIA Research(NVIDIA研究)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 15 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10315 2025-10-14 cs.LG cs.AI 57%

A Vision-Language Pre-training Model-Guided Approach for Mitigating Backdoor Attacks in Federated Learning

Keke Gai, Dongjue Wang, Jing Yu, Liehuang Zhu, Qi Wu

机构 * School of Cyberspace Science and Technology, Beijing Institute of Technology(信息空间科学与技术学院,北京理工大学) School of Information Engineering, Minzu University of China(信息工程学院,民族大学) Australian Institute of Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09867 2025-10-14 cs.CV 57%

Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation

Zhi Chen, Xin Yu, Xiaohui Tao, Yan Li, Zi Huang

机构 * University of Southern Queensland(昆士兰南方大学) University of Queensland(昆士兰大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to the journal Pattern Recognition in 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09699 2025-10-14 cs.CR cs.AI 57%

VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands

Aofan Liu, Lulu Tang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04479 2025-10-14 cs.CV 57%

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

Nonghai Zhang, Zeyu Zhang, Jiazi Wang, Yang Zhao, Hao Tang

机构 * Peking University(北京大学) Beijing Jiaotong University(北京交通大学) La Trobe University(拉特罗布大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08114 2025-10-14 cs.CV 57%

RATLIP: Generative Adversarial CLIP Text-to-Image Synthesis Based on Recurrent Affine Transformations

Chengde Lin, Xijun Lu, Guangxi Chen

机构 * School of Artificial Intelligence, Guangxi Colleges and Universities Key Laboratory of AI Algorithm Engineering(人工智能学院、广西 Colleges and Universities AI 算法工程重点实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by 2024 IEEE International Conference on Systems, Man, and Cybernetics(SMC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08659 2025-10-13 cs.LG cs.AI 57%

Provably Robust Adaptation for Language-Empowered Foundation Models

Yuni Lai, Xiaoyu Xue, Linghui Shen, Yulun Wu, Gaolei Li, Song Guo, Kai Zhou, Bin Xiao

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) College of Systems and Engineering, National University of Defense Technology(国防科技大学系统工程学院) School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院) Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08238 2025-10-10 cs.AI 57%

Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness

Jiyang Qiu, Xinbei Ma, Yunqing Xu, Zhuosheng Zhang, Hai Zhao

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07098 2025-10-09 cs.CL 57%

TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription

Guo Yutong, Wanying Wang, Yue Wu, Zichen Miao, Haoyu Wang

机构 * Johns Hopkins University(约翰霍普金斯大学) Purdue University(普渡大学) University at Albany(阿尔巴尼大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22688 2025-10-08 cs.CV 57%

Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization

Xu Jia

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17132 2025-10-08 cs.AI 57%

Applications of Large Models in Medicine

YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu, Yuxin Jia, Lujia Ge, Jing-min Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22229 2025-10-07 cs.CV 57%

A Tale of Two Experts: Cooperative Learning for Source-Free Unsupervised Domain Adaptation

Jiaping Yu, Muli Yang, Jiapeng Ji, Jiexi Yan, Cheng Deng

机构 * School of Electronic Engineering(电子工程学院) Xidian University(西安电子科技大学) School of Computer Science(计算机科学学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03858 2025-10-07 cs.CV 57%

Cross-View Open-Vocabulary Object Detection in Aerial Imagery

Jyoti Kini, Rohit Gupta, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18857 2025-10-07 cs.CV cs.LG 57%

Probabilistic Language-Image Pre-Training

Sanghyuk Chun, Wonjae Kim, Song Park, Sangdoo Yun

机构 * NAVER AI Lab(NAVER AI实验室)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Code: https://github.com/naver-ai/prolip HuggingFace Hub: https://huggingface.co/collections/SanghyukChun/prolip-6712595dfc87fd8597350291 33 pages, 4.5 MB; LongProLIP paper: arXiv:2503.08048; Multiplicity paper for more background: arxiv.org:2505.19614; v4: fix typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02780 2025-10-06 cs.CV 57%

Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models

Prahitha Movva

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Journal ref COLM 2025: First Workshop on the Application of LLM Explainability to Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09047 2025-10-06 cs.CL 57%

Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs

Yaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan Belinkov

机构 * Technion – Israel Institute of Technology(技术学院 – 以色列理工学院) UC Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18035 2025-10-06 cs.CV 57%

Vehicle-Scene Interaction: A Text-Driven 3D Lidar Place Recognition Method for Autonomous Driving

Tianyi Shang, Zhenyu Li, Pengjie Xu, Zhaojun Deng

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01534 2025-10-06 cs.CV 57%

Toward a Holistic Evaluation of Robustness in CLIP Models

Weijie Tu, Weijian Deng, Tom Gedeon

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to IEEE TPAMI, extension of NeurIPS'23 work: A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)

详情

展开后加载摘要…

URL PDF HTML 收藏