arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2311.02482 2023-11-07 cs.SD cs.AI cs.LG eess.AS 62%

Generalized zero-shot audio-to-intent classification

Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi, Sravan Bodapati, Srikanth Ronanki

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.10629 2023-11-06 eess.AS cs.CV cs.SD 62%

CSLNSpeech: solving extended speech separation problem with the help of Chinese sign language

Jiasong Wu, Xuan Li, Taotao Li, Fanman Meng, Youyong Kong, Guanyu Yang, Lotfi Senhadji, Huazhong Shu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments 13 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00867 2023-11-03 eess.AS cs.CL 62%

Automatic Disfluency Detection from Untranscribed Speech

Amrit Romana, Kazuhito Koishida, Emily Mower Provost

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06851 2023-10-12 cs.CV cs.AI cs.GR 62%

BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer

Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost, Takaaki Shiratori, Junichi Yamagishi, Taku Komura

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08021 2023-10-11 cs.CV cs.AI 62%

Vision-based Analysis of Driver Activity and Driving Performance Under the Influence of Alcohol

Ross Greer, Akshay Gopalkrishnan, Sumega Mandadi, Pujitha Gunaratne, Mohan M. Trivedi, Thomas D. Marcotte

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Withdrawn at the request of industry research collaborators, per contract agreement

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05203 2023-10-10 eess.AS cs.CL cs.LG cs.SD eess.SP 62%

A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023

Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta, Wen-Chin Huang, Tomoki Toda

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.CL、eess.AS

Comments Accepted to ASRU 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16308 2023-09-29 cs.MM cs.SD eess.AS 62%

Audio Visual Speaker Localization from EgoCentric Views

Jinzheng Zhao, Yong Xu, Xinyuan Qian, Wenwu Wang

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16058 2023-09-29 cs.LG cs.CL cs.CV 62%

AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Tushar Nagarajan, Matt Smith, Shashank Jain, Chun-Fu Yeh, Prakash Murugesan, Peyman Heidari, Yue Liu, Kavya Srinet, Babak Damavandi, Anuj Kumar

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13347 2023-09-26 cs.CL cs.SD eess.AS 62%

My Science Tutor (MyST) -- A Large Corpus of Children's Conversational Speech

Sameer S. Pradhan, Ronald A. Cole, Wayne H. Ward

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10537 2023-09-20 eess.AS cs.MM cs.SD 62%

FoleyGen: Visually-Guided Audio Generation

Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, Vikas Chandra

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06572 2023-09-14 eess.AS cs.CL cs.SD 62%

Addressing the Blind Spots in Spoken Language Processing

Amit Moryossef

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15898 2023-09-12 cs.SD cs.AI eess.AS 62%

UniBriVL: Robust Universal Representation and Generation of Audio Driven Diffusion Models

Sen Fang, Bowen Gao, Yangjian Wu, Teik Toe Teoh

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Voice-Text fusion input; The first work of audio driven diffusion model. arXiv admin note: text overlap with arXiv:2303.04585

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03295 2023-09-08 cs.CV cs.AI 62%

Comparative Analysis of Deep-Fake Algorithms

Nikhil Sontakke, Sejal Utekar, Shivansh Rastogi, Shriraj Sonawane

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 4 figures, 2 tables, Published with International Journal of Computer Science Trends and Technology (IJCST)

Journal ref International Journal of Computer Science Trends and Technology (IJCST) V11(4): Page(109-115) Jul - Aug 2023. ISSN: 2347-8578

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13736 2023-08-29 cs.SD cs.AI cs.HC eess.AS 62%

A Comprehensive Survey for Evaluation Methodologies of AI-Generated Music

Zeyu Xiong, Weitao Wang, Jing Yu, Yue Lin, Ziyan Wang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11329 2023-08-22 cs.CV cs.SD eess.AS 62%

Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation

Ziyang Chen, Shengyi Qian, Andrew Owens

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments ICCV 2023. Project site: https://ificl.github.io/SLfM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08125 2023-08-17 cs.SD cs.CL cs.HC eess.AS 62%

Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals

Running Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. Ngai

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted by Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (ACM IMWUT/UbiComp 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15377 2023-08-16 eess.AS cs.CV cs.LG cs.NE cs.SD 62%

Whose Emotion Matters? Speaking Activity Localisation without Prior Knowledge

Hugo Carneiro, Cornelius Weber, Stefan Wermter

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments 17 pages, 8 figures, 7 tables, Published in Neurocomputing

Journal ref Neurocomputing (2023); Volume 545; 126271

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04517 2023-08-10 cs.SD cs.CL eess.AS 62%

Capturing Spectral and Long-term Contextual Information for Speech Emotion Recognition Using Deep Learning Techniques

Samiul Islam, Md. Maksudul Haque, Abu Jobayer Md. Sadat

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments the research paper is still in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03504 2023-08-03 cs.CV cs.SD eess.AS 62%

Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis

Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu, Chen Zhang, Xiang Yin, Zejun Ma, Zhou Zhao

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments Accepted by ICML 2023 Workshop, 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13953 2023-07-27 cs.CV cs.SD eess.AS 62%

The Hidden Dance of Phonemes and Visage: Unveiling the Enigmatic Link between Phonemes and Facial Features

Liao Qu, Xianwei Zou, Xiang Li, Yandong Wen, Rita Singh, Bhiksha Raj

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09635 2023-07-25 cs.SD cs.LG cs.MM eess.AS eess.SP 62%

CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models

Hao-Wen Dong, Xiaoyu Liu, Jordi Pons, Gautam Bhattacharya, Santiago Pascual, Joan Serrà, Taylor Berg-Kirkpatrick, Julian McAuley

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM、eess.AS

Comments Accepted by WASPAA 2023. Demo: https://salu133445.github.io/clipsonic/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11450 2023-07-24 eess.AS cs.CL 62%

Topic Identification For Spontaneous Speech: Enriching Audio Features With Embedded Linguistic Information

Dejan Porjazovski, Tamás Grósz, Mikko Kurimo

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to EUSIPCO 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17203 2023-07-03 cs.SD cs.CV cs.LG eess.AS 62%

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Simian Luo, Chuanhao Yan, Chenxu Hu, Hang Zhao

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15808 2023-06-29 cs.MM cs.SD eess.AS eess.SP 62%

Classification of Infant Sleep/Wake States: Cross-Attention among Large Scale Pretrained Transformer Networks using Audio, ECG, and IMU Data

Kai Chieh Chang, Mark Hasegawa-Johnson, Nancy L. McElwain, Bashima Islam

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM、eess.AS

Comments Preprint for APSIPA2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09553 2023-06-28 cs.CL cs.SD eess.AS 62%

Mu$^{2}$SLAM: Multitask, Multilingual Speech and Language Models

Yong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey, Ankur Bapna

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14564 2023-06-23 cs.SD cs.AI eess.AS 62%

Exploring Self-supervised Pre-trained ASR Models For Dysarthric and Elderly Speech Recognition

Shujie Hu, Xurong Xie, Zengrui Jin, Mengzhe Geng, Yi Wang, Mingyu Cui, Jiajun Deng, Xunying Liu, Helen Meng

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments accepted by ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05040 2023-06-22 cs.CL cs.LG cs.SD eess.AS 62%

PATCorrect: Non-autoregressive Phoneme-augmented Transformer for ASR Error Correction

Ziji Zhang, Zhehui Wang, Rajesh Kamma, Sharanya Eswaran, Narayanan Sadagopan

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted camera-ready version for INTERSPEECH 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12020 2023-06-22 eess.AS cs.CL cs.SD 62%

Visual-Aware Text-to-Speech

Mohan Zhou, Yalong Bai, Wei Zhang, Ting Yao, Tiejun Zhao, Tao Mei

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments accepted as oral and top 3% paper by ICASSP 2023

Journal ref ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2023, 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09944 2023-06-19 cs.SD cs.CV cs.GR eess.AS 62%

RealImpact: A Dataset of Impact Sound Fields for Real Objects

Samuel Clarke, Ruohan Gao, Mason Wang, Mark Rau, Julia Xu, Jui-Hsien Wang, Doug L. James, Jiajun Wu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments CVPR 2023 (Highlight). Project page: https://samuelpclarke.com/realimpact/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06410 2023-06-13 cs.CL cs.CV 62%

OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

Xize Cheng, Tao Jin, Linjun Li, Wang Lin, Xinyu Duan, Zhou Zhao

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.CL

Comments Accepted to ACL2023 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏