arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-14 至 2025-08-14 共收录 66 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2410.19925 2025-08-14 cs.CL cs.CV cs.LG 84%

Improving Multimodal Large Language Models Using Continual Learning

Shikhar Srivastava, Md Yousuf Harun, Robik Shrestha, Christopher Kanan

机构 * University of Rochester(罗切斯特大学) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments CoLLAs 2025 and Scalable Continual Learning for Lifelong Foundation Models, NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09218 2025-08-14 cs.CV cs.AI 81%

Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity

Zuoou Li, Weitong Zhang, Jingyuan Wang, Shuyuan Zhang, Wenjia Bai, Bernhard Kainz, Mengyun Qiao

专题命中 图文多模态 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09199 2025-08-14 cs.CV cs.AI cs.CL 75%

$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation

Jucheng Hu, Suorong Yang, Dongzhan Zhou

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06220 2025-08-14 cs.CL cs.AI 62%

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

Keummin Ka, Junhyeong Park, Jaehyun Jeon, Youngjae Yu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06685 2025-08-14 cs.MM cs.CV 62%

Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding

Dawei Huang, Qing Li, Chuan Yan, Zebang Cheng, Zihao Han, Yurong Huang, Xiang Li, Bin Li, Xiaohui Wang, Zheng Lian, Zhi-Qi Cheng, Xiaojiang Peng

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06564 2025-08-14 cs.CV 57%

Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

Guanyu Hu, Dimitrios Kollias, Xinyu Yang

机构 * Xi'an Jiaotong University(西安交通大学) Queen Mary University of London(伦敦女王玛丽大学) Center for Multimodal AI(多模态人工智能中心) Digital Environment Research Institute(数字环境研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments accepted for publication at ACM Multimedia (ACM MM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22398 2025-08-14 cs.CV 57%

On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations

Jordan Vice, Naveed Akhtar, Yansong Gao, Richard Hartley, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Melbourne(墨尔本大学) Australian National University(澳大利亚国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Keywords: Vision-Language Models, Frequency-Domain Perturbations, Adversarial Robustness, Image Authenticity, Reliability

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17609 2025-08-14 cs.AI 57%

Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving

Zixian Guo, Ming Liu, Qilong Wang, Zhilong Ji, Jinfeng Bai, Lei Zhang, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨理工大学) The Hong Kong Polytechnic University(香港理工大学) Tianjin University(天津大学) Tomorrow Advancing Life(明天进步生命) Pazhou Lab, Guangzhou(广州琶洲实验室)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2508.09913 2025-08-14 cs.CV 79%

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

Yachao Liang, Min Yu, Gang Li, Jianguo Jiang, Boquan Li, Feng Yu, Ning Zhang, Xiang Meng, Weiqing Huang

机构 * Institute of Information Engineering(信息工程研究所) Chinese Academy of Sciences(中国科学院) School of Cyber Security University of Chinese Academy of Sciences(中国科学院网络安全学院) Deakin University(德肯大学) Harbin Engineering University(哈尔滨工程大学) Institute of Computing Technology Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of Forensic Science Ministry of Public Security(公安部刑事科学技术研究所)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2024

Journal ref Advances in Neural Information Processing Systems, Volume 37, Pages 86124-86144, Year 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09702 2025-08-14 eess.AS cs.SD 79%

$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation

Boyu Zhu, Cheng Gong, Muyang Wu, Ruihao Jing, Fan Liu, Xiaolei Zhang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAl), China Telecom(人工智能研究院(TeleAl),中国电信) School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09402 2025-08-14 cs.HC 78%

Realtime Multimodal Emotion Estimation using Behavioral and Neurophysiological Data

Von Ralph Dane Marquez Herbuela, Yukie Nagai

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12867 2025-08-14 eess.AS cs.AI cs.CL 67%

EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Guanrou Yang, Chen Yang, Qian Chen, Ziyang Ma, Wenxi Chen, Wen Wang, Tianrui Wang, Yifan Yang, Zhikang Niu, Wenrui Liu, Fan Yu, Zhihao Du, Zhifu Gao, ShiLiang Zhang, Xie Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Tongyi Speech Lab(通义语音实验室) Tianjin University(天津大学) Zhejiang University(浙江大学) Shanghai Jiao Tong University, Shanghai Innovation Institute(上海交通大学上海创新研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10417 2025-08-14 cs.CL cs.AI cs.SD eess.AS 67%

Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance

Abdelrahman A. Ali, Aya E. Fouda, Radwa J. Hanafy, Mohammed E. Fouda

机构 * Compumacy for Artificial Intelligence solutions(人工智能解决方案公司) Department of Behavioural Health- Saint Elizabeths Hospital(行为健康部门-圣伊丽莎白医院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2508.09789 2025-08-14 cs.IR cs.CV 83%

Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations

Marco De Nadai, Andreas Damianou, Mounia Lalmas

机构 * Spotify Denmark(Spotify丹麦分公司) Spotify United Kingdom(Spotify英国分公司)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09362 2025-08-14 cs.CV cs.AI cs.LG 81%

FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition

Md. Milon Islam, Md Rezwanul Haque, S M Taslim Uddin Raju, Fakhri Karray

机构 * University of Waterloo(滑铁卢大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for the IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, Hawaii, USA. 1st MSLR Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12341 2025-08-14 cs.MM 74%

Multimodal LLM-based Query Paraphrasing for Video Search

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Sheng-Hua Zhong, Xiong-Yong Wei, Qing Li

专题命中 视频多模态 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09857 2025-08-14 cs.CV 57%

OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better

Yupeng Zhou, Zhen Li, Ziheng Ouyang, Yuming Chen, Ruoyi Du, Daquan Zhou, Bin Fu, Yihao Liu, Peng Gao, Ming-Ming Cheng, Qibin Hou

机构 * VCIP, School of Computer Science, Nankai University(VCIP,计算机科学学院,南开大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18923 2025-08-14 cs.CV 57%

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Chen Wang, Jiahua Dong, Wangbo Yu, Ge Zhang, Jun Song, Xiang Li, Bo Zheng, Ian Reid, Xiaodan Liang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09262 2025-08-14 cs.CV cs.LG 57%

Harnessing Input-Adaptive Inference for Efficient VLN

Dongwoo Kang, Akhil Perincherry, Zachary Coalson, Aiden Gabriel, Stefan Lee, Sanghyun Hong

机构 * Oregon State University(俄勒冈州立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025 [Poster]

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2508.09170 2025-08-14 cs.LG cs.AI cs.CV cs.IR 81%

Multimodal RAG Enhanced Visual Description

Amit Kumar Jaiswal, Haiming Liu, Ingo Frommholz

机构 * Indian Institute of Technology (BHU)(印度理工学院(BHU)) University of Southampton(南安普顿大学) Modul University Vienna(维也纳应用科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM CIKM 2025. 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14700 2025-08-14 cs.CV cs.AI 81%

Depth-Guided Self-Supervised Human Keypoint Detection via Cross-Modal Distillation

Aman Anand, Elyas Rashno, Amir Eskandari, Farhana Zulkernine

机构 * Queen’s University Ontario Canada(皇后大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 11 篇

2504.19860 2025-08-14 cs.CV 85%

CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback

Chenhan Jiang, Yihan Zeng, Dit-Yan Yeung

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09945 2025-08-14 cs.CL cs.AI cs.CV 82%

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

Lingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li, Dongdong Zhang, Furu Wei

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09987 2025-08-14 cs.CV cs.AI cs.CL 67%

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Junyan Ye, Dongzhi Jiang, Zihao Wang, Leqi Zhu, Zhenghao Hu, Zilong Huang, Jun He, Zhiyuan Yan, Jinghua Yu, Hongsheng Li, Conghui He, Weijia Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sun Yat-sen University(中山大学) CUHK MMLab(香港中文大学多模态实验室) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09834 2025-08-14 cs.CL cs.AI cs.CV 67%

Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

Weigao Sun, Jiaxi Hu, Yucheng Zhou, Jusen Du, Disen Lan, Kexin Wang, Tong Zhu, Xiaoye Qu, Yu Zhang, Xiaoyu Mo, Daizong Liu, Yuxuan Liang, Wenliang Chen, Guoqi Li, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) HKUST (GZ)(香港科技大学) University of Macau(澳门大学) Institute of Automation Chinese Academy of Sciences(中国科学院自动化研究所) Soochow University(苏州大学) KTH Royal Institute of Technology(皇家理工学院) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Survey, 82 pages, GitHub: https://github.com/weigao266/Awesome-Efficient-Arch

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05296 2025-08-14 cs.AI cs.HC cs.SD eess.AS 62%

Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation

Joonwoo Kwon, Heehwan Wang, Jinwoo Lee, Sooyoung Kim, Shinjae Yoo, Yuewei Lin, Jiook Cha

机构 * Michigan State University(密歇根州立大学) Seoul National University(首尔国立大学) Rutgers University(罗格斯大学) Brookhaven National Laboratory(布鲁赫斯国家实验室)

专题命中 多模态生成 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted at the ACM MM 2025 - The 1st CogMAEC Workshop (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09461 2025-08-14 cs.CV cs.AI 62%

Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy

Hao Yu, Rupayan Mallick, Margrit Betke, Sarah Adel Bargal

机构 * Boston University(波士顿大学) Georgetown University(乔治城大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09225 2025-08-14 eess.IV cs.AI cs.CV 62%

AMRG: Extend Vision Language Models for Automatic Mammography Report Generation

Nak-Jun Sung, Donghyun Lee, Bo Hwa Choi, Chae Jung Park

机构 * Research Institute, National Cancer Center Korea(韩国国家癌症中心研究所) Department of Radiology, National Cancer Center Korea(韩国国家癌症中心放射科)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09177 2025-08-14 eess.IV cs.AI cs.CV 62%

Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation

Xuanru Zhou, Cheng Li, Shuqiang Wang, Ye Li, Tao Tan, Hairong Zheng, Shanshan Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01657 2025-08-14 cs.IR cs.CV 57%

RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation

Run Ling, Wenji Wang, Yuting Liu, Guibing Guo, Haowei Liu, Jian Lu, Quanwei Zhang, Yexing Xu, Shuo Lu, Yun Wang, Yihua Shao, Zhanjie Zhang, Ao Ma, Linying Jiang, Xingwei Wang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏