arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-12 至 2025-11-12 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 6 篇

2511.08031 2025-11-12 cs.CV cs.AI 86%

Multi-modal Deepfake Detection and Localization with FPN-Transformer

Chende Zheng, Ruiqi Suo, Zhoulin Ji, Jingyi Deng, Fangbin Yi, Chenhao Lin, Chao Shen

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13630 2025-11-12 cs.CV 85%

AVAR-Net: A Lightweight Audio-Visual Anomaly Recognition Framework with a Benchmark Dataset

Amjid Ali, Zulfiqar Ahmad Khan, Altaf Hussain, Muhammad Munsif, Adnan Hussain, Sung Wook Baik

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments I would like to request the withdrawal of my paper . The reason for this request is that I am currently working on additional experiments and analyses, which will lead to updates in the results section. Once these updates are complete, I will resubmit the revised version. Thank you for your understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10069 2025-11-12 cs.DC cs.LG 82%

ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism

Zedong Liu, Shenggan Cheng, Guangming Tan, Yang You, Dingwen Tao

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted at NeurIPS 2025 Oral (Thirty-Ninth Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19281 2025-11-12 cs.RO 82%

Audio-Visual Traffic Light State Detection for Urban Robots

Sagar Gupta, Akansel Cosgun

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);multi-modal(abstract)

Comments Submitted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13979 2025-11-12 cs.CL 79%

Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection

Maya Srikanth, Run Chen, Julia Hirschberg

机构 * Columbia University(哥伦比亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments To appear in Findings of IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09205 2025-11-12 cs.MM cs.CL cs.IR cs.SD eess.AS 67%

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

机构 * University of Rochester(罗切斯特大学) Microsoft Research(微软研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted at EUSIPCO 2025 - 5 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏