arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-26 至 2025-08-26 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 9 篇

2508.17502 2025-08-26 cs.CV 83%

Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice

Hugo Bohy, Minh Tran, Kevin El Haddad, Thierry Dutoit, Mohammad Soleymani

机构 * Numediart Institute, ISIA Lab, University of Mons(Numediart研究院、ISIA实验室、蒙斯大学) Institute for Creative Technologies, University of Southern California(创意技术研究所、南加州大学)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments 5 pages, 3 figures, IEEE FG 2024 conference

Journal ref 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11804 2025-08-26 cs.CR cs.AI cs.LG 83%

Adversarial Illusions in Multi-Modal Embeddings

Tingwei Zhang, Rishi Jha, Eugene Bagdasaryan, Vitaly Shmatikov

机构 * Cornell University(康奈尔大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Cornell Tech(康奈尔科技)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments In USENIX Security'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16606 2025-08-26 cs.HC cs.AI cs.LG 79%

Multimodal Appearance based Gaze-Controlled Virtual Keyboard with Synchronous Asynchronous Interaction for Low-Resource Settings

Yogesh Kumar Meena, Manish Salvi

机构 * Human-AI Interaction (HAIx) Lab, IIT Gandhinagar(人机交互(HAIx)实验室,印度加尔各答理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15632 2025-08-26 cs.SD 78%

ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification

Bochao Sun, Dong Wang, ZhanLong Yang, Jun Yang, Han Yin

机构 * School of Marine Science and Technology, Northwestern Polytechnical University, Xi’an, China(海洋科学与技术学院,西北工业大学,西安,中国) School of Automation, Northwestern Polytechnical University, Xi’an, China(自动化学院,西北工业大学,西安,中国) School of Electrical Engineering, KAIST, Daejeon, Republic of Korea(电气工程学院,韩国成均馆大学,大田,韩国)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01263 2025-08-26 cs.MM cs.CV cs.SD eess.AS 67%

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

Gaoxiang Cong, Liang Li, Jiadong Pan, Zhedong Zhang, Amin Beheshti, Anton van den Hengel, Yuankai Qi, Qingming Huang

机构 * Institute of Computing Technology, Chinese Academy of Sciences \& University of Chinese Academy of Sciences Beijing China Institute of Computing Technology, Chinese Academy of Sciences Beijing China Hangzhou Dianzi University Hangzhou China Macquarie University Sydney Australia University of Adelaide Adelaide Australia University of Chinese Academy of Sciences Beijing China Institute of Computing Technology, Chinese Academy of Sciences \& University of Chinese Academy of Sciences Institute of Computing Technology, Chinese Academy of Sciences Hangzhou Dianzi University Macquarie University University of Adelaide University of Chinese Academy of Sciences

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13237 2025-08-26 eess.AS cs.CL cs.SD 62%

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information

Chih-Kai Yang, Neo Ho, Yen-Ting Piao, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to Interspeech 2025 (Oral). Update acknowledgement in this version. Project page: https://github.com/ckyang1124/SAKURA

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16789 2025-08-26 cs.CV cs.CL 62%

Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation

Jungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young Kim

机构 * Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学) Graduate School of Information, Yonsei University(信息研究生院,延世大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07354 2025-08-26 cs.SD cs.IR cs.LG cs.MM eess.AS 62%

SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset

Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Mubarak Shah

机构 * SimulaMet OsloMet Forzasys University of Central Florida(佛罗里达中央大学) University of Liège(列日大学) KAUST(科威特大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17124 2025-08-26 cs.HC cs.SY eess.SY 50%

Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks

Ryan Ghamandi, Yahya Hmaiti, Mykola Maslych, Ravi Kiran Kattoju, Joseph J. LaViola

专题命中 音频语音多模态 :multimodal(abstract)

Comments To be submitted in a future conference, this is the author version pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏