arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-13 至 2025-10-13 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2506.18369 2025-10-13 cs.CV 83%

RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models

Yeongtak Oh, Dohyun Chung, Juhyeon Shin, Sangha Park, Johan Barthelemy, Jisoo Mok, Sungroh Yoon

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09358 2025-10-13 cs.CV 79%

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, Shaodong Chen, Yingyi Zhang, Chao Feng, Jiao Ran

机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11034 2025-10-13 cs.LG cs.AI cs.CL 62%

CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models

Aneesh Komanduri, Karuna Bhaila, Xintao Wu

机构 * Department of Electrical Engineering and Computer Science University of Arkansas(电气工程与计算机科学系 奎萨克大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025 Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03355 2025-10-13 cs.LG cs.AI cs.CV 62%

Robustness in Both Domains: CLIP Needs a Robust Text Encoder

Elias Abad Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu, Matthias Hein, Volkan Cevher

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08659 2025-10-13 cs.LG cs.AI 57%

Provably Robust Adaptation for Language-Empowered Foundation Models

Yuni Lai, Xiaoyu Xue, Linghui Shen, Yulun Wu, Gaolei Li, Song Guo, Kai Zhou, Bin Xiao

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) College of Systems and Engineering, National University of Defense Technology(国防科技大学系统工程学院) School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院) Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 1 篇

2509.22425 2025-10-13 cs.SD 78%

From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation

Ke Xue, Rongfei Fan, Lixin, Dawei Zhao, Chao Zhu, Han Hu

机构 * School of Cyberspace Science and Technology(网络空间科学与技术学院) Beijing Institute of Technology(北京理工大学) Sun Yat-sen University(中山大学) Qilu University of Technology(齐鲁工业大学) Shandong Computer Science Center(山东计算机科学中心) School of Information and Electronics(信息电子学院)

专题命中 音频语音多模态 :audio-visual(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2510.09266 2025-10-13 cs.CL 79%

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang, Ruida Liu, Yuming Yang, Xin Xiao, Xiao Sun, Haoyang Zeng, Changzai Pan, Yidan Zhang, Jiang Zhong, Peijin Wang, Yingchao Feng

机构 * Chongqing University(重庆大学) Independent Researcher(独立研究者) University of the Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06509 2025-10-13 cs.CV 74%

From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding

Shih-Yao Lin, Sibendu Paul, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02362 2025-10-13 cs.CV cs.AI 62%

A Comprehensive Survey of Mamba Architectures for Medical Image Analysis: Classification, Segmentation, Restoration and Beyond

Shubhi Bansal, Sreeharish A, Madhava Prasath J, Manikandan S, Sreekanth Madisetty, Mohammad Zia Ur Rehman, Chandravardhan Singh Raghaw, Gaurav Duggal, Nagendra Kumar

机构 * Indian Institute of Technology Indore, India R.M.D. Engineering College, Kavaraipettai, India Jio Platforms Limited, India Birla Institute of Technology \& Science Pilani, India

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09200 2025-10-13 cs.CV cs.AI cs.HC 62%

Towards Safer and Understandable Driver Intention Prediction

Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C V Jawahar

机构 * IIIT Hyderabad(海得拉巴印度理工学院) Politecnico di Torino(托里诺理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03991 2025-10-13 cs.CV 57%

Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges

Hao Xu, Arbind Agrahari Baniya, Sam Well, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal

机构 * School of Information Technology, Deakin University(德肯大学信息科技学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2508.04379 2025-10-13 cs.CV cs.LG 79%

VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones

Lefei Shen, Mouxiang Chen, Xu Liu, Han Fu, Xiaoxue Ren, Jianling Sun, Zhuo Li, Chenghao Liu

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) State Street Technology (Zhejiang) Ltd.(State Street Technology(浙江)有限公司) Salesforce Research Asia(Salesforce亚洲研究)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20582 2025-10-13 cs.IR 78%

SUMMA: A Multimodal Large Language Model for Advertisement Summarization

Weitao Jia, Shuo Yin, Zhoufutu Wen, Han Wang, Zehui Dai, Kun Zhang, Zhenyu Li, Tao Zeng, Xiaohui Lv

专题命中 跨模态检索 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11101 2025-10-13 cs.CV cs.LG 74%

A Survey on Self-supervised Contrastive Learning for Multimodal Text-Image Analysis

Asifullah Khan, Laiba Asmatullah, Anza Malik, Shahzaib Khan, Hamna Asif

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments 38 pages, 8 figures, survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09586 2025-10-13 cs.CV 57%

Vision Language Models: A Survey of 26K Papers

Fengming Lin

机构 * School of Computer Science, The University of Manchester, Manchester, UK(曼彻斯特大学计算机科学学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments VLM/LLM Learning Notes

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08578 2025-10-13 cs.MA cs.AI cs.HC 57%

AgenticAD: A Specialized Multiagent System Framework for Holistic Alzheimer Disease Management

Adib Bazgir, Amir Habibdoust, Xing Song, Yuwen Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 6 篇

2509.21787 2025-10-13 cs.CV cs.CL 81%

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal

机构 * MIT Media Lab, USA(麻省理工学院媒体实验室) Stanford University, USA(斯坦福大学) Amazon GenAI, USA(亚马逊生成人工智能) University of South Carolina, USA(南卡罗来纳大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Defactify 3 workshop at AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24840 2025-10-13 cs.LG cs.CE 78%

Cell2Text: Multimodal LLM for Generating Single-Cell Descriptions from RNA-Seq Data

Oussama Kharouiche, Aris Markogiannakis, Xiao Fei, Michail Chatzianastasis, Michalis Vazirgiannis

机构 * École Polytechnique, IP Paris(巴黎理工学院,IP巴黎)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09254 2025-10-13 cs.RO cs.AI 57%

Obstacle Avoidance using Dynamic Movement Primitives and Reinforcement Learning

Dominik Urbaniak, Alejandro Agostini, Pol Ramon, Jan Rosell, Raúl Suárez, Michael Suppa

机构 * Institute of Industrial and Control Engineering, Universitat Politècnica de Catalunya(工业与控制工程研究所,巴塞罗那理工大学) Department of Computer Science, University of Innsbruck(计算机科学系,因斯布鲁克大学) Roboception GmbH(Roboception公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08625 2025-10-13 cs.CV 57%

Adjusting Initial Noise to Mitigate Memorization in Text-to-Image Diffusion Models

Hyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong, Jungwoo Lee

机构 * Seoul National University(首尔国立大学) CSE, Konkuk University(计算机科学与工程系,konkuk大学) Hodoo AI Labs(Hodoo AI实验室)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09134 2025-10-13 cs.SE cs.ET 50%

A Semantic Framework for Patient Digital Twins in Chronic Care

Amal Elgammal, Bernd J. Krämer, Michael P. Papazoglou, Mira Raheem

专题命中 多模态生成 :multimodal(abstract)

Comments This manuscript is currently under review at Software and Systems Modeling (SoSyM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08912 2025-10-13 cs.HC 50%

Beyond Words: Infusing Conversational Agents with Human-like Typing Behaviors

Jijie Zhou, Yuhan Hu

专题命中 多模态生成 :multimodal(abstract)

Comments Author's version of a paper published at CUI '24 (ACM Conversational User Interfaces 2024)

Journal ref CUI '24: Proceedings of the ACM Conversational User Interfaces 2024, July 8-10, 2024, Luxembourg, Luxembourg. ACM, New York, NY, USA, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 13 篇

2510.08783 2025-10-13 cs.HC cs.AI 86%

MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces

Reuben A. Luera, Ryan Rossi, Franck Dernoncourt, Samyadeep Basu, Sungchul Kim, Subhojyoti Mukherjee, Puneet Mathur, Ruiyi Zhang, Jihyung Kil, Nedim Lipka, Seunghyun Yoon, Jiuxiang Gu, Zichao Wang, Cindy Xiong Bearfield, Branislav Kveton

机构 * University of California, Berkeley(加州大学伯克利分校) Adobe Research(Adobe研究) Georgia Institute of Technology(佐治亚理工学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08608 2025-10-13 cs.CL cs.AI 84%

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Agency for Science, Technology and Research, Singapore(新加坡科技研究局) Indian Institute of Technology Delhi(印度理工学院德里分校) Alibaba DAMO Academy(阿里巴巴达摩院) Microsoft Research Asia(微软亚洲研究院) Shanghai University of Finance and Economics(上海财经大学) Inner Mongolia University(内蒙古大学) Kyoto University(京都大学) Jiangxi Normal University(江西师范大学) Korea University(韩国大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09230 2025-10-13 cs.CV cs.AI cs.CL cs.LG 82%

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) Peking University People’s Hospital(北京大学人民医院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15175 2025-10-13 cs.RO 82%

SHeRLoc: Synchronized Heterogeneous Radar Place Recognition for Cross-Modal Localization

Hanjun Kim, Minwoo Jung, Wooseong Yang, Ayoung Kim

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract)

Comments 9 pages, 9 figures, accepted to RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08964 2025-10-13 cs.CV cs.CL 81%

Unleashing Perception-Time Scaling to Multimodal Reasoning Models

Yifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou, Ruipu Luo, Can Zhang, Zhentao He, Yufei Zhan, Wayne Xin Zhao, Minghui Qiu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理重点实验室) ByteDance(字节跳动) University of California, San Diego(加州大学圣地亚哥分校) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08928 2025-10-13 cs.AI 79%

LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

Yushuo Zheng, Zicheng Zhang, Xiongkuo Min, Huiyu Duan, Guangtao Zhai

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09407 2025-10-13 q-fin.GN cs.LG 71%

A Multimodal Approach to SME Credit Scoring Integrating Transaction and Ownership Networks

Sahab Zandi, Kamesh Korangi, Juan C. Moreno-Paredes, María Óskarsdóttir, Christophe Mues, Cristián Bravo

机构 * Department of Statistical and Actuarial Sciences, Western University(西敏大学统计与精算科学系) University of Southampton Business School(南安普顿大学商学院) Centre for Operational Research, Management Sciences and Information Systems(运营研究、管理科学与信息系统中心) School of Mathematical Sciences, University of Southampton(南安普顿大学数学科学学院) Department of Computer Science, Reykjavík University(雷克雅未克大学计算机科学系)

专题命中 多模态评测 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08978 2025-10-13 cs.CV 70%

HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images

Zichuan Wang, Bo Peng, Songlin Yang, Zhenchen Tang, Jing Dong

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏