arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2305.00359 2023-05-31 eess.AS 57%

A Review of Deep Learning Techniques for Speech Processing

Ambuj Mehrish, Navonil Majumder, Rishabh Bhardwaj, Rada Mihalcea, Soujanya Poria

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17137 2023-05-30 cs.AI cs.LG 57%

Integrating Generative Artificial Intelligence in Intelligent Vehicle Systems

Lukas Stappen, Jeremy Dillmann, Serena Striegel, Hans-Jörg Vögel, Nicolas Flores-Herr, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05534 2023-05-10 cs.CV cs.LG 57%

Integrating Holistic and Local Information to Estimate Emotional Reaction Intensity

Yini Fang, Liang Wu, Frederic Jumelle, Bertram Shi

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments This paper will be published in CVPRW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01241 2023-05-09 cs.HC cs.GR cs.LG cs.SD eess.AS 57%

AQ-GT: a Temporally Aligned and Quantized GRU-Transformer for Co-Speech Gesture Synthesis

Hendric Voß, Stefan Kopp

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00537 2023-05-02 cs.MM cs.CY cs.LG 57%

Interpretability of Machine Learning: Recent Advances and Future Prospects

Lei Gao, Ling Guan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

Comments IEEE Multimedia (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10871 2023-05-02 cs.LG cs.CL cs.SI 57%

Qualitative Analysis of a Graph Transformer Approach to Addressing Hate Speech: Adapting to Dynamically Changing Content

Liam Hebert, Hong Yi Chen, Robin Cohen, Lukasz Golab

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments Accepted at AAAI 2023 AI for Social Good

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06786 2023-04-28 eess.AS cs.SD 57%

The future of hearing aid technology

Volker Hohmann

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06246 2023-04-06 cs.LG cs.CV cs.SD 57%

Jointly Learning Visual and Auditory Speech Representations from Raw Data

Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis, Maja Pantic

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments ICLR 2023. Code: https://github.com/ahaliassos/raven

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16856 2023-03-30 cs.CV cs.GR 57%

Robust Dancer: Long-term 3D Dance Synthesis Using Unpaired Data

Bin Feng, Tenglong Ao, Zequn Liu, Wei Ju, Libin Liu, Ming Zhang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Preliminary video demo: https://youtu.be/gJbxG9QlcUU

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12379 2023-03-23 cs.CV 57%

VMCML: Video and Music Matching via Cross-Modality Lifting

Yi-Shan Lee, Wei-Cheng Tseng, Fu-En Wang, Min Sun

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10667 2023-03-21 cs.SD eess.AS 57%

Audio-Text Models Do Not Yet Leverage Natural Language

Ho-Hsiang Wu, Oriol Nieto, Juan Pablo Bello, Justin Salamon

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Copyright 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03569 2023-03-21 cs.HC cs.AI 57%

Human in the Loop for Machine Creativity

Neo Christopher Chung

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 9th AAAI Conference on Human Computation and Crowdsourcing (HCOMP 2021), Blue Sky Ideas track

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13523 2023-03-15 cs.SD eess.AS 57%

VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting

Ao Zhang, He Wang, Pengcheng Guo, Yihui Fu, Lei Xie, Yingying Gao, Shilei Zhang, Junlan Feng

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 5 pages. Accepted at ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06357 2023-03-14 cs.CV 57%

CASP-Net: Rethinking Video Saliency Prediction from an Audio-VisualConsistency Perceptual Perspective

Junwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang, Yufei Zha, Guangtao Zhai

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 10 pages, 7 figures, CVPR2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01906 2023-03-14 cs.RO cs.AI cs.HC 57%

A trained humanoid robot can perform human-like crossmodal social attention and conflict resolution

Di Fu, Fares Abawi, Hugo Carneiro, Matthias Kerzel, Ziwei Chen, Erik Strahl, Xun Liu, Stefan Wermter

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments accepted for publication in the International Journal of Social Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00109 2023-03-10 eess.AS cs.SD 57%

ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding Inpainting

Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11960 2023-02-28 cs.SD cs.LG eess.AS 57%

Disentangled Feature Learning for Real-Time Neural Speech Coding

Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments ICASSP 2023 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09817 2023-02-24 cs.LG cs.CV 57%

Explainable Human-centered Traits from Head Motion and Facial Expression Dynamics

Surbhi Madan, Monika Gahalawat, Tanaya Guha, Roland Goecke, Ramanathan Subramanian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08137 2023-02-17 cs.SD cs.LG eess.AS 57%

ACE-VC: Adaptive and Controllable Voice Conversion using Explicitly Disentangled Self-supervised Speech Representations

Shehzeen Hussain, Paarth Neekhara, Jocelyn Huang, Jason Li, Boris Ginsburg

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Published as a conference paper at ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12808 2023-01-31 eess.AS 57%

Real-Time Acoustic Perception for Automotive Applications

Jun Yin, Stefano Damiano, Marian Verhelst, Toon van Waterschoot, Andre Guntoro

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10062 2023-01-25 cs.CV cs.IR cs.LG 57%

Proceedings of the 1st International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Jan Hajič, Alexander Pacha

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza, Jan Hajič jr. and Alexander Pacha

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06690 2023-01-18 cs.CV 57%

Audio2Gestures: Generating Diverse Gestures from Audio

Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Linchao Bao, Zhenyu He

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2108.06720

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.08918 2023-01-18 cs.LG cs.AI 57%

Neuro-SERKET: Development of Integrative Cognitive System through the Composition of Deep Probabilistic Generative Models

Tadahiro Taniguchi, Tomoaki Nakamura, Masahiro Suzuki, Ryo Kuniyasu, Kaede Hayashi, Akira Taniguchi, Takato Horii, Takayuki Nagai

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments New Gener. Comput. (2020)

Journal ref New Generation Computing, 2020, volume 38, 23--48

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05063 2022-12-13 cs.HC cs.AI 57%

The RPM3D project: 3D Kinematics for Remote Patient Monitoring

Alicia Fornés, Asma Bensalah, Cristina Carmona-Duarte, Jialuo Chen, Miguel A. Ferrer, Andreas Fischer, Josep Lladós, Cristina Martín, Eloy Opisso, Réjean Plamondon, Anna Scius-Bertrand, Josep Maria Tormos

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00380 2022-12-02 cs.CV cs.IR cs.LG 57%

Proceedings of the 2nd International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Alexander Pacha

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza and Alexander Pacha

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00378 2022-12-02 cs.CV cs.IR cs.LG 57%

Proceedings of the 3rd International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Alexander Pacha

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza and Alexander Pacha

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14739 2022-12-02 cs.AI cs.NE 57%

A Case for Business Process-Specific Foundation Models

Yara Rizk, Praveen Venkateswaran, Vatche Isahagian, Vinod Muthusamy

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15058 2022-11-29 cs.CV 57%

Mix and Localize: Localizing Sound Sources in Mixtures

Xixi Hu, Ziyang Chen, Andrew Owens

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.05437 2022-11-29 cs.RO cs.AI cs.LG cs.NE 57%

Autonomous Racing using a Hybrid Imitation-Reinforcement Learning Architecture

Chinmay Vilas Samak, Tanmay Vilas Samak, Sivanathan Kandhasamy

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04356 2022-11-23 cs.SD cs.LG eess.AS 57%

A Comparative Study of Self-supervised Speech Representation Based Voice Conversion

Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Tomoki Toda

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Accepted to IEEE Journal of Selected Topics in Signal Processing. arXiv admin note: substantial text overlap with arXiv:2110.06280

详情

展开后加载摘要…

URL PDF HTML 收藏