arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-27 至 2025-10-27 共收录 56 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2510.21346 2025-10-27 cs.CV cs.AI 84%

CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments

Lemin Liu, Fangchao Hu, Honghua Jiang, Yaru Chen, Limin Liu, Yongliang Qiao

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21093 2025-10-27 cs.AI 79%

MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning

Siyong Chen, Jinbo Wen, Jiawen Kang, Tenghui Huang, Xumin Huang, Yuanjia Su, Hudan Pan, Zishao Zhong, Dusit Niyato, Shengli Xie, Dong In Kim

机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21606 2025-10-27 cs.CV 70%

Modest-Align: Data-Efficient Alignment for Vision-Language Models

Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Tianyi Zhou, Mingkun Xu, Zuozhu Liu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙大天使align智能医疗研发中心) Centre for Frontier AI Research (CFAR)(前沿人工智能研究中心) Agency for Science, Technology and Research (A*STAR)(科技研究局) Institute of High Performance Computing (IHPC)(高性能计算研究所)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21311 2025-10-27 cs.CV 70%

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

Lu Zhang, Jiazuo Yu, Haomiao Xiong, Ping Hu, Yunzhi Zhuge, Huchuan Lu, You He

机构 * Dalian University of Technology(大连理工大学) University of Electronic Science and Technology of China(电子科技大学) Tsinghua Shenzhen International Graduate School(清华大学深圳国际graduate school)

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21323 2025-10-27 cs.CV cs.LG 57%

VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set

Shufan Shen, Junshu Sun, Qingming Huang, Shuhui Wang

机构 * Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS(智能信息处理重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21083 2025-10-27 cs.CV 57%

Knowledge-Driven Vision-Language Model for Plexus Detection in Hirschsprung's Disease

Youssef Megahed, Atallah Madi, Dina El Demellawy, Adrian D. C. Chan

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted into the ICAAI 2025 - The 9th International Conference on Advances in Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 10 篇

2504.16102 2025-10-27 cs.CV cs.RO 84%

HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues

Xiwen Li, Xiaoya Tang, Tolga Tasdizen

机构 * Scientific Computing and Imaging Institute(科学计算与成像研究所) Department of Electrical and Computer Engineering(电气与计算机工程系) University of Utah(犹他大学)

专题命中 音频语音多模态 :cross-modal(title);audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12995 2025-10-27 eess.AS cs.SD 83%

Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs

Xinlu He, Swayambhu Nath Ray, Harish Mallidi, Jia-Hong Huang, Ashwin Bellur, Chander Chandak, M. Maruf, Venkatesh Ravichandran

机构 * Worcester Polytechnic Institute(沃斯特理工大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11217 2025-10-27 cs.SD cs.AI cs.CV cs.MM eess.AS 77%

Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization

Yanhao Jia, Ji Xie, S Jivaganesh, Hao Li, Xu Wu, Mengmi Zhang

机构 * Nanyang Technological University, Singapore(南洋理工大学) Peking University, China(北京大学) Shenzhen University, China(深圳大学)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments NeurIPS 2025, Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01957 2025-10-27 cs.CV cs.SD eess.AS 74%

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Chaoyou Fu, Haojia Lin, Xiong Wang, Yi-Fan Zhang, Yunhang Shen, Xiaoyu Liu, Haoyu Cao, Zuwei Long, Heting Gao, Ke Li, Long Ma, Xiawu Zheng, Rongrong Ji, Xing Sun, Caifeng Shan, Ran He

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Tencent Youtu Lab(腾讯优图实验室) XMU(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 音频语音多模态 :MLLM(abstract,comments);multimodal(abstract);分类 cs.CV、eess.AS

Comments NeurIPS 2025 Spotlight, Code 2.4K Stars: https://github.com/VITA-MLLM/VITA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22088 2025-10-27 cs.SD cs.CL eess.AS 62%

Visual Cues Support Robust Turn-taking Prediction in Noise

Sam O'Connor Russell, Naomi Harte

机构 * ADAPT Centre, School of Engineering(ADAPT中心,工程学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to INTERSPEECH 2025, 10.21437/Interspeech.2025-668

Journal ref Proc. of INTERSPEECH 2025, 1073--1077, 10.21437/Interspeech.2025-668

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20850 2025-10-27 eess.AS cs.CL cs.SD 62%

Can large audio language models understand child stuttering speech? speech summarization, and source separation

Chibuzor Okocha, Maya Bakri, Christan Grant

机构 * Department of Computer Science, University of Florida, Gainesville, FL, USA(佛罗里达大学计算机科学系) Department of Computer Science, Lebanese American University(黎巴嫩美国大学计算机科学系)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 7 pages, 1 Figure, 8 tables, Under review ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21581 2025-10-27 cs.CV cs.SD 57%

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer, Julian Parker, Zach Evans

机构 * Stability AI

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://stability-ai.github.io/foleycontrol.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21043 2025-10-27 cs.CL cs.RO 57%

Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction

Sam O'Connor Russell, Naomi Harte

机构 * ADAPT Centre, School of Engineering, Trinity College Dublin(ADAPT中心、工程学院、都柏林信任学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted to ACL 2025, Findings of the Association for Computational Linguistics

Journal ref In Findings of the Association for Computational Linguistics: ACL 2025, pages 209--221, Vienna, Austria. Association for Computational Linguistics, 10.18653/v1/2025.findings-acl.12

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18972 2025-10-27 cs.LG cs.AI 57%

Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques

David Ortiz-Perez, Manuel Benavent-Lledo, Jose Garcia-Rodriguez, David Tomás, M. Flores Vizcaya-Moreno

机构 * Dept. of Computer Science and Technology(计算机科学与技术系) University of Alicante(阿尔瓦登特大学) Unit of Clinical Nursing Research(临床护理研究单位) Faculty of Health Sciences(健康科学学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Journal ref Applied Soft Computing, Vol. 184, 2025, Article 113787

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20867 2025-10-27 cs.LG cs.AI 57%

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

Jiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey, Prashanth Gurunath Shivakumar, Ivan Bulyko, Ankur Gandhe, Ge Liu, Yile Gu

机构 * Amazon(亚马逊) Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(计算机与数据科学学院,伊利诺伊大学厄巴纳-香槟分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 49 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 12 篇

2510.21445 2025-10-27 cs.CL cs.AI cs.CV cs.LG 82%

REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring

Thanh Cong Ho, Farah Kharrat, Abderrazek Abid, Fakhri Karray

机构 * 2 Department of Electrical Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email 3 College of Computer Information Sciences Prince Sultan University Email

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21406 2025-10-27 cs.CV 79%

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

Yue Feng, Jinwei Hu, Qijia Lu, Jiawei Niu, Li Tan, Shuo Yuan, Ziyi Yan, Yizhen Jia, Qingzhi He, Shiping Ge, Ethan Q. Chen, Wentong Li, Limin Wang, Jie Qin

机构 * MoE Key Laboratory of Brain-Machine Intelligence Technology, College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(脑机智能技术MoE实验室,人工智能学院,南京航空航天大学) Nanjing University(南京大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03340 2025-10-27 cs.CV 79%

Seeing the Arrow of Time in Large Multimodal Models

Zihui Xue, Mi Luo, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025, Project website: https://vision.cs.utexas.edu/projects/SeeAoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20952 2025-10-27 cs.LG 78%

LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting

Sungjun Cho, Changho Shin, Suenggwan Jo, Xinya Yan, Shourjo Aditya Chaudhuri, Frederic Sala

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 视频多模态 :multimodal(title,abstract)

Comments 15 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18438 2025-10-27 cs.CR 71%

DeepTx: Real-Time Transaction Risk Analysis via Multi-Modal Features and LLM Reasoning

Yixuan Liu, Xinlei Li, Yi Li

专题命中 视频多模态 :multi-modal(title)

Comments Accepted to ASE'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03674 2025-10-27 cs.CV cs.AI 62%

Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression

Mengshi Qi, Hao Ye, Jiaxuan Peng, Huadong Ma

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20189 2025-10-27 cs.CV 57%

SPAN: Continuous Modeling of Suspicion Progression for Temporal Intention Localization

Xinyi Hu, Yuran Wang, Ruixu Zhang, Yue Li, Wenxuan Liu, Zheng Wang

机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science(多媒体软件国家工程研究中心、人工智能研究院、计算机科学学院) Hubei Key Laboratory of Multimedia and Network Communication Engineering(多媒体与网络通信工程湖北省重点实验室) School of Mathematical Sciences, Peking University(北京大学数学科学学院) Tsinghua University(清华大学) School of Computer Science, Peking University(北京大学计算机科学学院) State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08974 2025-10-27 cs.CV 57%

Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering

Elman Ghazaei, Erchan Aptoula

机构 * Faculty of Engineering and Natural Sciences (VPALab)(工程与自然科学学院(VPALab)) Sabanci University(萨班奇大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18812 2025-10-27 cs.CV 57%

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang, Xingjun Ma, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21107 2025-10-27 cs.LG cs.AI cs.RO 57%

ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs

Yunuo Zhang, Baiting Luo, Ayan Mukhopadhyay, Gabor Karsai, Abhishek Dubey

机构 * Vanderbilt University(范德比大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments Proceeding of the 39th Conference on Neural Information Processing Systems (NeurIPS'25). Code would be available at https://github.com/scope-lab-vu/ESCORT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20951 2025-10-27 cs.CV 57%

Generative Point Tracking with Flow Matching

Mattie Tesfaldet, Adam W. Harley, Konstantinos G. Derpanis, Derek Nowrouzezahrai, Christopher Pal

机构 * McGill University(麦吉尔大学) Mila Stanford University(斯坦福大学) York University(约克大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Project page: https://mtesfaldet.net/genpt_projpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15745 2025-10-27 eess.IV cs.LG 50%

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

专题命中 视频多模态 :multimodal(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2505.11293 2025-10-27 cs.CV 57%

Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining

Raghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou, Wenhu Chen, Bhuwan Dhingra

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 17 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 3 篇

2412.06771 2025-10-27 cs.AI cs.CV cs.LG 62%

Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty

Meera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim, Zi Wang

机构 * Google DeepMind(谷歌DeepMind)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Journal ref International Conference on Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏