arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2407.02264 2025-04-01 cs.CV cs.SD eess.AS 62%

SOAF: Scene Occlusion-aware Neural Acoustic Field

Huiyu Gao, Jiahao Ma, David Ahmedt-Aristizabal, Chuong Nguyen, Miaomiao Liu

机构 * Australian National University(澳大利亚国立大学) CSIRO Data61

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22275 2025-03-31 eess.AS cs.AI cs.SD 62%

Make Some Noise: Towards LLM audio reasoning and generation using sound tokens

Shivam Mehta, Nebojsa Jojic, Hannes Gamper

机构 * KTH Royal Institute of Technology(瑞典皇家理工学院) Microsoft Research(微软研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 2 figures, Accepted at ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08295 2025-03-26 cs.CL cs.SD eess.AS 62%

SpeechVerse: A Large-scale Generalizable Audio Language Model

Nilaksh Das, Saket Dingliwal, Srikanth Ronanki, Rohit Paturi, Zhaocheng Huang, Prashant Mathur, Jie Yuan, Dhanush Bekal, Xing Niu, Sai Muralidhar Jayanthi, Xilai Li, Karel Mundnich, Monica Sunkara, Sravan Bodapati, Sundararajan Srinivasan, Kyu J Han, Katrin Kirchhoff

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Single Column, 13 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18042 2025-03-21 cs.CV cs.CL 62%

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Kai Chen, Yunhao Gou, Runhui Huang, Zhili Liu, Daxin Tan, Jing Xu, Chunwei Wang, Yi Zhu, Yihan Zeng, Kuo Yang, Dingdong Wang, Kun Xiang, Haoyuan Li, Haoli Bai, Jianhua Han, Xiaohui Li, Weike Jin, Nian Xie, Yu Zhang, James T. Kwok, Hengshuang Zhao, Xiaodan Liang, Dit-Yan Yeung, Xiao Chen, Zhenguo Li, Wei Zhang, Qun Liu, Jun Yao, Lanqing Hong, Lu Hou, Hang Xu

机构 * Hong Kong University of Science and Technology(香港科技大学) The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚方舟实验室) The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) Southern University of Science and Technology(南方科技大学)

专题命中 音频语音多模态 :omni-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by CVPR 2025. Project Page: https://emova-ollm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15338 2025-03-20 eess.AS cs.CL cs.SD 62%

Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context

Junyi Ao, Dekun Chen, Xiaohai Tian, Wenjie Feng, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, Zhizheng Wu

机构 * School of Data Science, Shenzhen Research Institute of Big Data, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院、深圳大数据研究院) Bytedance(字节跳动)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14185 2025-03-19 cs.CL cs.SD eess.AS 62%

AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation

Wuwei Huang, Dexin Wang, Deyi Xiong

机构 * College of Intelligence and Computing, Tianjin University(天津大学智能与计算学部)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments ACL 2021 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08920 2025-03-18 cs.SD cs.AI eess.AS 62%

AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis

Swapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu

机构 * University of Surrey(萨里大学) Imperial College London(帝国理工学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17202 2025-03-13 cs.SD cs.CL eess.AS 62%

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

Chen Chen, Yuchen Hu, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang, Chao-Han Huck Yang, Eng Siong Chng

机构 * Nanyang Technological University(南洋理工大学) NVIDIA(英伟达) Tsinghua University(清华大学) Johns Hopkins University(约翰斯·霍普金斯大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08540 2025-03-12 cs.SD cs.AI eess.AS 62%

Mellow: a small audio language model for reasoning

Soham Deshmukh, Satvik Dixit, Rita Singh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Checkpoint and dataset available at: https://github.com/soham97/mellow

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19924 2025-02-28 cs.SD cs.AI eess.AS 62%

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

Weihao wu, Zhiwei Lin, Yixuan Zhou, Jingbei Li, Rui Niu, Qinghua Wu, Songjun Cao, Long Ma, Zhiyong Wu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) The Chinese University of Hong Kong(香港中文大学) Tencent Youtu Lab(腾讯优图实验室) StepFun(阶跃星辰)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00385 2025-02-27 cs.CL cs.AI 62%

The Impact of Persona-based Political Perspectives on Hateful Content Detection

Stefano Civelli, Pietro Bernardelle, Gianluca Demartini

机构 * The University of Queensland(昆士兰大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Companion Proceedings of the ACM Web Conference 2025 (WWW Companion'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16857 2025-02-25 cs.CL cs.AI 62%

Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models

Avinash Trivedi, Sangeetha Sivanesan

机构 * National Institute of Technology, Tiruchirapalli(印度特里奇拉巴利国家理工学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments AAAI-25 DEFACTIFY 4.0 Workshop AI generated text detection (1st Rank)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16137 2025-02-25 cs.CL cs.AI 62%

Chain-of-Description: What I can understand, I can put into words

Jiaxin Guo, Daimeng Wei, Zongyao Li, Hengchao Shang, Yuanchang Luo, Hao Yang

机构 * Huawei Translation Services Center(华为翻译服务中心)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14892 2025-02-24 cs.CV cs.AI 62%

EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild

Junhyeok Kim, Min Soo Kim, Jiwan Chung, Jungbin Cho, Jisoo Kim, Sungwoong Kim, Gyeongbo Sim, Youngjae Yu

机构 * Yonsei University(延世大学) NCSOFT Corporation(NCsoft公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments NAACL 2025 Findings. Project page at https://jun297.github.io/EgoSpeak/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13983 2025-02-21 eess.AS cs.AI 62%

Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders

Seungbae Kim, Daeun Lee, Brielle Stark, Jinyoung Han

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03739 2025-02-21 cs.CL cs.AI 62%

Grammar Induction from Visual, Speech and Text

Yu Zhao, Hao Fei, Shengqiong Wu, Meishan Zhang, Min Zhang, Tat-seng Chua

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16746 2025-02-18 cs.LG cs.AI cs.CL 62%

The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources

Shayne Longpre, Stella Biderman, Alon Albalak, Hailey Schoelkopf, Daniel McDuff, Sayash Kapoor, Kevin Klyman, Kyle Lo, Gabriel Ilharco, Nay San, Maribeth Rauh, Aviya Skowron, Bertie Vidgen, Laura Weidinger, Arvind Narayanan, Victor Sanh, David Adelani, Percy Liang, Rishi Bommasani, Peter Henderson, Sasha Luccioni, Yacine Jernite, Luca Soldaini

机构 * MIT(麻省理工学院) EleutherAI UCSB(加利福尼亚大学圣巴巴拉分校) SynthLabs Princeton University(普林斯顿大学) Stanford University(斯坦福大学) Harvard University(哈佛大学) Allen Institute for AI(艾伦人工智能研究所) Google DeepMind(谷歌DeepMind) ML Commons Contextual AI HuggingFace University College London(伦敦大学学院)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09940 2025-02-17 cs.CL cs.SD eess.AS 62%

A Preliminary Exploration with GPT-4o Voice Mode

Yu-Xiang Lin, Chih-Kai Yang, Wei-Chih Chen, Chen-An Li, Chien-yu Huang, Xuanjun Chen, Hung-yi Lee

机构 * National Taiwan University(台湾大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07538 2025-02-14 cs.MM cs.SD eess.AS 62%

Visual-based spatial audio generation system for multi-speaker environments

Xiaojing Liu, Ogulcan Gurelli, Yan Wang, Joshua Reiss

机构 * Queen Mary University of London(伦敦玛丽女王大学) Xidian University(西安电子科技大学)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06922 2025-02-12 cs.SD cs.AI cs.CL cs.LG 62%

Synthetic Audio Helps for Cognitive State Tasks

Adil Soubki, John Murzaku, Peter Zeng, Owen Rambow

机构 * Stony Brook University(石溪大学) Department of Computer Science(计算机科学系) Department of Linguistics(语言学系) Institute for Advanced Computational Science(高级计算科学研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments John Murzaku and Adil Soubki contributed equally to this work

Journal ref NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03128 2025-02-06 cs.SD cs.AI cs.LG eess.AS eess.SP 62%

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Yuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang, Huan Liao, Zhizheng Wu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01046 2025-02-04 cs.SD cs.CV eess.AS 62%

Emotional Face-to-Speech

Jiaxin Ye, Boyuan Cao, Hongming Shan

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18801 2025-02-03 cs.CV cs.AI 62%

Every Image Listens, Every Image Dances: Music-Driven Image Animation

Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Peng Zhang, Pawel Polak

机构 * Stony Brook University(石溪大学) Bytedance(字节跳动) Apple(苹果公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11170 2025-01-22 cs.CL cs.AI cs.LG 62%

AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis

Alireza Ghahramani Kure, Mahshid Dehghani, Mohammad Mahdi Abootorabi, Nona Ghazizadeh, Seyed Arshan Dalili, Ehsaneddin Asgari

机构 * Sharif University of Technology(谢里夫理工大学) Qatar Computing Research Institute(卡塔尔计算研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10937 2025-01-22 cs.CL cs.SD eess.AS 62%

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data

Jingran Xie, Shun Lei, Yue Yu, Yang Xiang, Hui Wang, Xixin Wu, Zhiyong Wu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Pengcheng Laboratory(鹏城实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09104 2025-01-22 cs.SD cs.AI eess.AS 62%

A Non-autoregressive Model for Joint STT and TTS

Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis Lastras

机构 * The Ohio State University(俄亥俄州立大学) IBM Research(IBM研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18836 2025-01-20 cs.SD cs.AI eess.AS 62%

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI

Neil Shah, Ayan Kashyap, Shirish Karande, Vineet Gandhi

机构 * IIIT Hyderabad(印度国际信息技术学院海得拉巴分校) TCS Research(塔塔咨询服务公司研究院)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted at IEEE ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09818 2025-01-17 cs.CL cs.AI 62%

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

Yingxu He, Zhuohan Liu, Shuo Sun, Bin Wang, Wenyu Zhang, Xunlong Zou, Nancy F. Chen, Ai Ti Aw

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments https://huggingface.co/MERaLiON/MERaLiON-AudioLLM-Whisper-SEA-LION

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01020 2025-01-14 cs.CV cs.SD eess.AS 62%

A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio

Xavier Juanola, Gloria Haro, Magdalena Fuentes

机构 * Universitat Pompeu Fabra(庞培法布拉大学) MARL-IDM, New York University(纽约大学MARL-IDM)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments Accepted in ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20619 2024-12-31 cs.LG cs.MM cs.SD eess.AS 62%

Audiopedia: Audio QA with Knowledge

Abhirama Subramanyam Penamakuri, Kiran Chhatre, Akshat Jain

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏