arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2212.06246 2023-04-06 cs.LG cs.CV cs.SD 57%

Jointly Learning Visual and Auditory Speech Representations from Raw Data

Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis, Maja Pantic

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments ICLR 2023. Code: https://github.com/ahaliassos/raven

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16856 2023-03-30 cs.CV cs.GR 57%

Robust Dancer: Long-term 3D Dance Synthesis Using Unpaired Data

Bin Feng, Tenglong Ao, Zequn Liu, Wei Ju, Libin Liu, Ming Zhang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Preliminary video demo: https://youtu.be/gJbxG9QlcUU

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12379 2023-03-23 cs.CV 57%

VMCML: Video and Music Matching via Cross-Modality Lifting

Yi-Shan Lee, Wei-Cheng Tseng, Fu-En Wang, Min Sun

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10667 2023-03-21 cs.SD eess.AS 57%

Audio-Text Models Do Not Yet Leverage Natural Language

Ho-Hsiang Wu, Oriol Nieto, Juan Pablo Bello, Justin Salamon

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Copyright 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03569 2023-03-21 cs.HC cs.AI 57%

Human in the Loop for Machine Creativity

Neo Christopher Chung

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 9th AAAI Conference on Human Computation and Crowdsourcing (HCOMP 2021), Blue Sky Ideas track

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13523 2023-03-15 cs.SD eess.AS 57%

VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting

Ao Zhang, He Wang, Pengcheng Guo, Yihui Fu, Lei Xie, Yingying Gao, Shilei Zhang, Junlan Feng

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 5 pages. Accepted at ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06357 2023-03-14 cs.CV 57%

CASP-Net: Rethinking Video Saliency Prediction from an Audio-VisualConsistency Perceptual Perspective

Junwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang, Yufei Zha, Guangtao Zhai

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 10 pages, 7 figures, CVPR2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01906 2023-03-14 cs.RO cs.AI cs.HC 57%

A trained humanoid robot can perform human-like crossmodal social attention and conflict resolution

Di Fu, Fares Abawi, Hugo Carneiro, Matthias Kerzel, Ziwei Chen, Erik Strahl, Xun Liu, Stefan Wermter

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments accepted for publication in the International Journal of Social Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00109 2023-03-10 eess.AS cs.SD 57%

ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding Inpainting

Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11960 2023-02-28 cs.SD cs.LG eess.AS 57%

Disentangled Feature Learning for Real-Time Neural Speech Coding

Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments ICASSP 2023 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09817 2023-02-24 cs.LG cs.CV 57%

Explainable Human-centered Traits from Head Motion and Facial Expression Dynamics

Surbhi Madan, Monika Gahalawat, Tanaya Guha, Roland Goecke, Ramanathan Subramanian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08137 2023-02-17 cs.SD cs.LG eess.AS 57%

ACE-VC: Adaptive and Controllable Voice Conversion using Explicitly Disentangled Self-supervised Speech Representations

Shehzeen Hussain, Paarth Neekhara, Jocelyn Huang, Jason Li, Boris Ginsburg

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Published as a conference paper at ICASSP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12808 2023-01-31 eess.AS 57%

Real-Time Acoustic Perception for Automotive Applications

Jun Yin, Stefano Damiano, Marian Verhelst, Toon van Waterschoot, Andre Guntoro

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10062 2023-01-25 cs.CV cs.IR cs.LG 57%

Proceedings of the 1st International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Jan Hajič, Alexander Pacha

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza, Jan Hajič jr. and Alexander Pacha

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06690 2023-01-18 cs.CV 57%

Audio2Gestures: Generating Diverse Gestures from Audio

Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Linchao Bao, Zhenyu He

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2108.06720

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.08918 2023-01-18 cs.LG cs.AI 57%

Neuro-SERKET: Development of Integrative Cognitive System through the Composition of Deep Probabilistic Generative Models

Tadahiro Taniguchi, Tomoaki Nakamura, Masahiro Suzuki, Ryo Kuniyasu, Kaede Hayashi, Akira Taniguchi, Takato Horii, Takayuki Nagai

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments New Gener. Comput. (2020)

Journal ref New Generation Computing, 2020, volume 38, 23--48

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05063 2022-12-13 cs.HC cs.AI 57%

The RPM3D project: 3D Kinematics for Remote Patient Monitoring

Alicia Fornés, Asma Bensalah, Cristina Carmona-Duarte, Jialuo Chen, Miguel A. Ferrer, Andreas Fischer, Josep Lladós, Cristina Martín, Eloy Opisso, Réjean Plamondon, Anna Scius-Bertrand, Josep Maria Tormos

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00380 2022-12-02 cs.CV cs.IR cs.LG 57%

Proceedings of the 2nd International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Alexander Pacha

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza and Alexander Pacha

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00378 2022-12-02 cs.CV cs.IR cs.LG 57%

Proceedings of the 3rd International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Alexander Pacha

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza and Alexander Pacha

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14739 2022-12-02 cs.AI cs.NE 57%

A Case for Business Process-Specific Foundation Models

Yara Rizk, Praveen Venkateswaran, Vatche Isahagian, Vinod Muthusamy

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15058 2022-11-29 cs.CV 57%

Mix and Localize: Localizing Sound Sources in Mixtures

Xixi Hu, Ziyang Chen, Andrew Owens

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.05437 2022-11-29 cs.RO cs.AI cs.LG cs.NE 57%

Autonomous Racing using a Hybrid Imitation-Reinforcement Learning Architecture

Chinmay Vilas Samak, Tanmay Vilas Samak, Sivanathan Kandhasamy

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.04356 2022-11-23 cs.SD cs.LG eess.AS 57%

A Comparative Study of Self-supervised Speech Representation Based Voice Conversion

Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Tomoki Toda

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Accepted to IEEE Journal of Selected Topics in Signal Processing. arXiv admin note: substantial text overlap with arXiv:2110.06280

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09376 2022-11-18 cs.SD cs.LG eess.AS 57%

Balanced Deep CCA for Bird Vocalization Detection

Sumit Kumar, B. Anshuman, Linus Ruettimann, Richard H. R. Hahnloser, Vipul Arora

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10608 2022-11-18 cs.CL 57%

Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora

Sara Papi, Alina Karakanta, Matteo Negri, Marco Turchi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Journal ref AACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14099 2022-11-18 cs.CL 57%

Climate and Weather: Inspecting Depression Detection via Emotion Recognition

Wen Wu, Mengyue Wu, Kai Yu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Journal ref ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6262-6266

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05446 2022-11-11 cs.SD cs.CR cs.LG eess.AS 57%

Privacy-Utility Balanced Voice De-Identification Using Adversarial Examples

Meng Chen, Li Lu, Jiadi Yu, Yingying Chen, Zhongjie Ba, Feng Lin, Kui Ren

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03511 2022-11-08 cs.CL 57%

End-to-End Evaluation of a Spoken Dialogue System for Learning Basic Mathematics

Eda Okur, Saurav Sahay, Roddy Fuentes Alba, Lama Nachman

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Proceedings of the 1st Workshop on Mathematical Natural Language Processing (MathNLP) at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03019 2022-11-08 cs.CV 57%

Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source Localization

Dennis Fedorishin, Deen Dayal Mohan, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02336 2022-11-07 cs.SD eess.AS 57%

Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts

Detai Xin, Sharath Adavanne, Federico Ang, Ashish Kulkarni, Shinnosuke Takamichi, Hiroshi Saruwatari

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏