arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-14 至 2025-10-14 共收录 120 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 17 篇

2507.04351 2025-10-14 cs.RO cs.AI 88%

MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection

Liman Wang, Hanyang Zhong, Tianyuan Wang, Shan Luo, Jihong Zhu

机构 * School of Physics, Engineering and Technology, University of York(物理、工程与技术学院,约克大学) Department of Engineering, King’s College London(工程学院,伦敦国王学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.AI

Comments Accepted to IEEE Robotics and Automation Letters (RAL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10560 2025-10-14 cs.CL cs.AI cs.CV 85%

BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices

Euhid Aman, Esteban Carlin, Hsing-Kuo Pao, Giovanni Beltrame, Ghaluh Indah Permata Sari, Yie-Tarng Chen

机构 * NTUST(国立台湾科技大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 6 pages, BabyLM Workshop, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09815 2025-10-14 cs.CV cs.AI 84%

Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning

Yufei Wang, Adriana Kovashka, Loretta Fernández, Marc N. Coutanche, Seth Wiener

机构 * University of Pittsburgh(匹兹堡大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to International Conference on Development and Learning (ICDL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10466 2025-10-14 cs.CV 83%

When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance

Jinjin Cao, Zhiyang Chen, Zijun Wang, Liyuan Ma, Weijian Luo, Guojun Qi

机构 * MAPLE Lab, Westlake University(西溪大学MAPLE实验室)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11115 2025-10-14 cs.CV cs.MM 81%

Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning

Hao Tang, Shengfeng He, Jing Qin

机构 * Centre for Smart Health, The Hong Kong Polytechnic University(香港理工大学智能健康中心) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算机与信息系统学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10342 2025-10-14 cs.CV 79%

Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis

Yu-Hsuan Lin

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 7 pages, 4 figures. Preprint submitted to arXiv in October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10117 2025-10-14 cs.AI 79%

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments EMNLP 2025 Wordplay (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10104 2025-10-14 cs.CV 79%

Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models

Minbin Huang, Runhui Huang, Chuanyang Zheng, Jingyao Li, Guoxuan Chen, Han Shi, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01738 2025-10-14 cs.CV 70%

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

Ming Dai, Wenxuan Cheng, Jiang-jiang Liu, Sen Yang, Wenxiao Cai, Yanpeng Sun, Wankou Yang

机构 * Southeast University(东南大学) Baidu VIS(百度VIS) Stanford University(斯坦福大学)

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22215 2025-10-14 cs.CV cs.AI cs.CL cs.LG 67%

Learning to Instruct for Visual Instruction Tuning

Zhihan Zhou, Feng Hong, Jiaan Luo, Jiangchao Yao, Dongsheng Li, Bo Han, Ya Zhang, Yanfeng Wang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学合作中元创新中心) Microsoft Research Asia(微软亚洲研究院) Hong Kong Baptist University(香港 Baptist 大学) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10052 2025-10-14 cs.CV cs.AI 62%

Think Twice to See More: Iterative Visual Reasoning in Medical VLMs

Kaitao Chen, Shaohao Rui, Yankai Jiang, Jiamin Wu, Qihao Zheng, Chunfeng Song, Xiaosong Wang, Mu Zhou, Mianxin Liu

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Rutgers University(罗格斯大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10962 2025-10-14 cs.LG cs.AI 57%

MC#: Mixture Compressor for Mixture-of-Experts Large Models

Wei Huang, Yue Liao, Yukang Chen, Jianhui Liu, Haoru Tan, Si Liu, Shiming Zhang, Shuicheng Yan, Xiaojuan Qi

机构 * Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系) School of Computing, National University of Singapore(新加坡国立大学计算机学院) NVIDIA Research(NVIDIA研究)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 15 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10315 2025-10-14 cs.LG cs.AI 57%

A Vision-Language Pre-training Model-Guided Approach for Mitigating Backdoor Attacks in Federated Learning

Keke Gai, Dongjue Wang, Jing Yu, Liehuang Zhu, Qi Wu

机构 * School of Cyberspace Science and Technology, Beijing Institute of Technology(信息空间科学与技术学院,北京理工大学) School of Information Engineering, Minzu University of China(信息工程学院,民族大学) Australian Institute of Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09867 2025-10-14 cs.CV 57%

Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation

Zhi Chen, Xin Yu, Xiaohui Tao, Yan Li, Zi Huang

机构 * University of Southern Queensland(昆士兰南方大学) University of Queensland(昆士兰大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to the journal Pattern Recognition in 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09699 2025-10-14 cs.CR cs.AI 57%

VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands

Aofan Liu, Lulu Tang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04479 2025-10-14 cs.CV 57%

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

Nonghai Zhang, Zeyu Zhang, Jiazi Wang, Yang Zhao, Hao Tang

机构 * Peking University(北京大学) Beijing Jiaotong University(北京交通大学) La Trobe University(拉特罗布大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08114 2025-10-14 cs.CV 57%

RATLIP: Generative Adversarial CLIP Text-to-Image Synthesis Based on Recurrent Affine Transformations

Chengde Lin, Xijun Lu, Guangxi Chen

机构 * School of Artificial Intelligence, Guangxi Colleges and Universities Key Laboratory of AI Algorithm Engineering(人工智能学院、广西 Colleges and Universities AI 算法工程重点实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by 2024 IEEE International Conference on Systems, Man, and Cybernetics(SMC)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 9 篇

2510.10051 2025-10-14 cs.CV 85%

Complementary and Contrastive Learning for Audio-Visual Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang, Pingping Zhang, Huchuan Lu

机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) School of Future Technology and the School of Artificial Intelligence, Dalian University of Technology(未来技术学院和人工智能学院,大连理工大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01181 2025-10-14 cs.AI cs.CV cs.MM cs.SD eess.AS 83%

Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning

Zhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song, Xun Yang

机构 * University of Science and Technology of China(科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ACM Multimedia 2025 Oral Code: https://github.com/ZhiyuanHan-Aaron/MoSEAR Project Page: https://zhiyuanhan-aaron.github.io/MoSEAR-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05295 2025-10-14 cs.SD cs.AI cs.MM 82%

AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement

M. Sajid, Deepanshu Gupta, Yash Modi, Sanskriti Jain, Harshith Jai Surya Ganji, A. Rahaman, Harshvardhan Choudhary, Nasir Saleem, Amir Hussain, M. Tanveer

机构 * Indian Institute of Technology Indore(印度理工学院印第安纳分校) School of Computing, Edinburgh Napier University(爱丁堡纳皮尔大学计算机学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、cs.MM

Journal ref INTERSPEECH 2025 - 4th COG-MHEAR Workshop on Audio-Visual Speech Enhancement (AVSEC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11596 2025-10-14 cs.HC 78%

GlobalizeEd: A Multimodal Translation System that Preserves Speaker Identity in Academic Lectures

Hoang-Son Vo, Karina Kolmogortseva, Ngumimi Karen Iyortsuun, Hong-Duyen Vo, Soo-Hyung Kim

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11330 2025-10-14 cs.SD cs.AI cs.CL cs.LG eess.AS 67%

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology, South Korea(韩国科学技术院) University of Seoul, South Korea(首尔大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments 5 pages. Submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10201 2025-10-14 cs.LG cs.AI cs.CL 62%

RLFR: Extending Reinforcement Learning for LLMs with Flow Environment

Jinghao Zhang, Naishan Zheng, Ruilin Li, Dongzhou Cheng, Zheming Liang, Feng Zhao, Jiaqi Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) ByteDance(字节跳动) Wuhan University(武汉大学) Southeast University(东南大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Project Website: https://jinghaoleven.github.io/RLFR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11454 2025-10-14 cs.SD cs.AI 57%

Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

Kuan-Yi Lee, Tsung-En Lin, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 9pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10329 2025-10-14 cs.CL 57%

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

Nam Luu, Ondřej Bojar

机构 * Charles University(查尔斯大学) Faculty of Mathematics and Physics(数学与物理系) Institute of Formal and Applied Linguistics(形式与应用语言学研究所)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10173 2025-10-14 cs.HC cs.CY cs.SD eess.AS 57%

Chord Colourizer: A Near Real-Time System for Visualizing Musical Key

Paul Haimes

机构 * Ritsumeikan University(立命馆大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Author copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 11 篇

2510.11110 2025-10-14 cs.LG cs.AI 79%

PhysioME: A Robust Multimodal Self-Supervised Framework for Physiological Signals with Missing Modalities

Cheol-Hui Lee, Hwa-Yeon Lee, Min-Kyung Jung, Dong-Joo Kim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10986 2025-10-14 cs.CV 79%

Mixup Helps Understanding Multimodal Video Better

Xiaoyu Ma, Ding Ding, Hao Chen

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其跨学科应用关键实验室(东南大学),教育部,中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08561 2025-10-14 cs.CV 79%

MultiCOIN: Multi-Modal COntrollable Video INbetweening

Maham Tanveer, Yang Zhou, Simon Niklaus, Ali Mahdavi Amiri, Hao Zhang, Krishna Kumar Singh, Nanxuan Zhao

机构 * Simon Fraser University(西蒙弗雷泽大学) Adobe Research(Adobe研究院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Project website: https://multicoinx.github.io/multicoin/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01116 2025-10-14 cs.RO cs.LG cs.SY eess.SY 78%

Scalable Multi-modal Model Predictive Control via Duality-based Interaction Predictions

Hansung Kim, Siddharth H. Nair, Francesco Borrelli

机构 * Model Predictive Control Laboratory, UC Berkeley(模型预测控制实验室,伯克利大学)

专题命中 视频多模态 :multi-modal(title,abstract)

Comments Accepted at IEEE Intelligent Vehicles Symposium 2024

详情

展开后加载摘要…

URL PDF HTML 收藏