arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-29 至 2025-08-29 共收录 47 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 5 篇

2508.20181 2025-08-29 cs.CV cs.AI cs.CL cs.MM 85%

Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

Alberto Compagnoni, Davide Caffagni, Nicholas Moratelli, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20691 2025-08-29 cs.CV cs.AI cs.CL cs.LG 85%

MobileCLIP2: Improving Multi-Modal Reinforced Training

Fartash Faghri, Pavan Kumar Anasosalu Vasu, Cem Koc, Vaishaal Shankar, Alexander Toshev, Oncel Tuzel, Hadi Pouransari

机构 * Apple(苹果公司)

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments TMLR August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20243 2025-08-29 cs.CV cs.LG 70%

Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification

Mutahar Safdar, Gentry Wood, Max Zimmermann, Guy Lamouche, Priti Wanjara, Yaoyao Fiona Zhao

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 46 pages, 33 figures, Submitted to Advanced Engineering Informatics, under revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10583 2025-08-29 cs.CV cs.CL 62%

Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models

Diogo Freitas, Brigt Håvardstun, Cèsar Ferri, Darío Garigliotti, Jan Arne Telle, José Hernández-Orallo

机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15576 2025-08-29 cs.CV cs.LG 57%

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

Xin Huang, Ruibin Li, Tong Jia, Wei Zheng, Ya Wang

机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 10 篇

2301.06375 2025-08-29 cs.MM cs.AI cs.CL cs.CV cs.LG cs.SD 85%

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyun Lee, Jun Hwan Ahn, Rae-Hong Park, Hyung-Min Park

机构 * 1 Department of Artificial Intelligence, Sogang University, Seoul 04107, Republic of Korea 2 Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea 3 Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA 4 Mindslab Inc., Gyeonggi-do 13493, Republic of Korea 5 ICT Convergence Disaster/Safety Research Institute, Sogang University, Seoul 04107, Republic of Korea

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16188 2025-08-29 cs.CL cs.CV cs.MM cs.SD eess.AS 85%

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello, Philipp Koehn, Xutai Ma

机构 * Johns Hopkins University(约翰霍普金斯大学) Meta AI Research(Meta AI 研究)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20546 2025-08-29 cs.MM cs.AI 84%

MM-HSD: Multi-Modal Hate Speech Detection in Videos

Berta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko, Andrea Cavallaro

机构 * EPFL(苏黎世联邦理工学院) Idiap Research Institute(日内瓦研究所)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted at ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20359 2025-08-29 cs.IR 82%

Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation

Shijia Wang, Tianpei Ouyang, Qiang Xiao, Dongjing Wang, Yintao Ren, Songpei Xu, Da Guo, Chuanjiang Luo

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20805 2025-08-29 cs.CL cs.AI cs.SD 81%

Exploring Machine Learning and Language Models for Multimodal Depression Detection

Javier Si Zhao Hong, Timothy Zoe Delaya, Sherwyn Chan Yin Kit, Pai Chet Ng, Xiaoxiao Miao

机构 * Singapore Institute of Technology(新加坡理工学院) Duke Kunshan University(杜克-昆山大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by APCIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20513 2025-08-29 cs.SD cs.MM 79%

MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening

Yongqi Shao, Binxin Mei, Cong Tan, Hong Huo, Tao Fang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20221 2025-08-29 cs.CV 79%

Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos

Mert Cokelek, Halit Ozsoy, Nevrez Imamoglu, Cagri Ozcinar, Inci Ayhan, Erkut Erdem, Aykut Erdem

机构 * Department of Computer Science and Engineering, Koç University(计算机科学与工程系,科克大学) Department of Psychology, Boğaziçi University(心理学系,博多伊大学) National Institute of Advanced Industrial Science and Technology (AIST), Intelligent Platforms Research Institute(国家先进工业科学与技术研究院(AIST),智能平台研究机构) Department of Psychology, Boğaziçi University University(心理学系,博多伊大学) Department of Computer Engineering, Hacettepe University(计算机工程系,哈切塞特佩大学) KUIS AI Center(KUIS人工智能中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted for publication in IEEE Transaction on Pattern Analysis and Machine Intelligence (IEEE TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20796 2025-08-29 cs.SD cs.AI 57%

Speech Emotion Recognition via Entropy-Aware Score Selection

ChenYi Chua, JunKai Wong, Chengxin Chen, Xiaoxiao Miao

机构 * Singapore Institute of Technology(新加坡理工学院) Institute Of Acoustics, Chinese Academy Of Sciences(中国科学院声学研究所) Duke Kunshan University(杜克大学昆山分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments The paper has been accepted by APCIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20782 2025-08-29 eess.AS 57%

A Solution of Ultra Wideband Based High-resolution and Lossless Audio Transmission

Fengyun Zhang

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17336 2025-08-29 cs.SD cs.AI 57%

Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework

Yunsik Kim, Yoonyoung Chung

机构 * Department of Electrical Engineering(电气工程系) Department of Semiconductor Engineering(半导体工程系) Center for Semiconductor Technology Convergence(半导体技术融合中心)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Journal ref Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2507.04651 2025-08-29 cs.IR 85%

FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation

Maolin Wang, Yutian Xiao, Binhao Wang, Sheng Zhang, Shanshan Ye, Wanyu Wang, Hongzhi Yin, Ruocheng Guo, Zenglin Xu

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)

Comments Accepted by KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20121 2025-08-29 cs.NE cs.SY eess.IV eess.SY 71%

Task-Aware Tuning of Time Constants in Spiking Neural Networks for Multimodal Classification

Chiu-Chang Cheng, Kapil Bhardwaj, Ya-Ning Chang, Sayani Majumdar, Chao-Hung Wang

专题命中 视频多模态 :multimodal(title)

Comments 25 Pages and 5 Figures and a supplementary discussion as well

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21070 2025-08-29 cs.CV cs.LG 57%

Dress&Dance: Dress up and Dance as You Like It - Technical Preview

Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) SpreeAI

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://immortalco.github.io/DressAndDance/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19958 2025-08-29 cs.RO 50%

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

Yiguo Fan, Pengxiang Ding, Shuanghao Bai, Xinyang Tong, Yuyang Zhu, Hongchao Lu, Fengqi Dai, Wei Zhao, Yang Liu, Siteng Huang, Zhaoxin Fan, Badong Chen, Donglin Wang

专题命中 视频多模态 :multimodal(abstract)

Comments Accepted to CoRL 2025; Github Page: https://long-vla.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2508.20188 2025-08-29 cs.CV cs.LG 83%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 78%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 跨模态检索 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11535 2025-08-29 stat.ML cs.LG cs.SY eess.SY stat.CO 50%

Canonical Bayesian Linear System Identification

Andrey Bryutkin, Matthew E. Levine, Iñigo Urteaga, Youssef Marzouk

机构 * Massachusetts Institute of Technology(麻省理工学院) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) Basis Research Institute(Basis研究机构) Eric and Wendy Schmidt Center(埃里克和温迪·施密特中心) BCAM (Basque Center for Applied Mathematics)(BCAM(巴斯克应用数学中心)) Ikerbasque (Basque Foundation for Science)(Ikerbasque(巴斯克科学基金会))

专题命中 跨模态检索 :multi-modal(abstract)

Comments 46 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 4 篇

2508.19320 2025-08-29 cs.CV cs.AI 81%

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

Ming Chen, Liyuan Cui, Wenyuan Zhang, Haoxian Zhang, Yan Zhou, Xiaohan Li, Songlin Tang, Jiwen Liu, Borui Liao, Hejia Chen, Xiaoqiang Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队) Zhejiang University(浙江大学) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Technical Report. Project Page: https://chenmingthu.github.io/milm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20379 2025-08-29 cs.CV 79%

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

Hyeonyu Kim, Seokhoon Jeong, Seonghee Han, Chanhyuk Choi, Taehwan Kim

机构 * MAUM AI Inc.(MAUM AI公司) Artificial Intelligence Graduate School UNIST(UNIST人工智能研究生学院)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20182 2025-08-29 cs.CV 70%

SDiFL: Stable Diffusion-Driven Framework for Image Forgery Localization

Yang Su, Shunquan Tan, Jiwu Huang

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19852 2025-08-29 cs.CV 57%

Ego-centric Predictive Model Conditioned on Hand Trajectories

Binjie Zhang, Mike Zheng Shou

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments Code: github.com/showlab/Ego-PM

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 7 篇

2508.20851 2025-08-29 cs.CV 83%

PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

Ye Zhang, Yu Zhou, Jingwen Qi, Yongbing Zhang, Simon Puettmann, Finn Wichmann, Larissa Pereira Ferreira, Lara Sichward, Julius Keyl, Sylvia Hartmann, Shuo Zhao, Hongxiao Wang, Xiaowei Xu, Jianxu Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14195 2025-08-29 cs.HC cs.CV 79%

A multimodal dataset for understanding the impact of mobile phones on remote online virtual education

Roberto Daza, Alvaro Becerra, Ruth Cobos, Julian Fierrez, Aythami Morales

机构 * Biometrics and Data Pattern Analytics Laboratory(生物信息与数据模式分析实验室) School of Engineering(工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Published in Scientific Data (Nature). GitHub repository of the dataset at: https://github.com/BiDAlab/IMPROVE

Journal ref Scientific Data (2025) 12:1332

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20583 2025-08-29 cs.CL cs.AI 62%

A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models

Soham Petkar, Hari Aakash K, Anirudh Vempati, Akshit Sinha, Ponnurangam Kumarauguru, Chirag Agarwal

机构 * plaksha.edu.in(普拉克斯哈大学) research.iiit.ac.in(IIIT研究机构)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20660 2025-08-29 eess.AS cs.SD 57%

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

Ruifan Deng, Yitian Gong, Qinghui Gao, Luozhijie Jin, Qinyuan Cheng, Zhaoye Fei, Shimin Li, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态评测 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏