arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2106.04133 2021-06-09 cs.SD cs.LG eess.AS 57%

Efficient Speech Emotion Recognition Using Multi-Scale CNN and Attention

Zixuan Peng, Yu Lu, Shengfeng Pan, Yunfeng Liu

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments First two authors contributed equally.Accepted by ICASSP 2021

Journal ref ICASSP,2021 pp. 3020-3024

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09939 2021-05-21 cs.CV 57%

Face, Body, Voice: Video Person-Clustering with Multiple Modalities

Andrew Brown, Vicky Kalogeiton, Andrew Zisserman

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.00650 2021-05-12 cs.RO cs.SD eess.AS 57%

Making Sense of Audio Vibration for Liquid Height Estimation in Robotic Pouring

Hongzhuo Liang, Shuang Li, Xiaojian Ma, Norman Hendrich, Timo Gerkmann, Fuchun Sun, Jianwei Zhang

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to IROS 2019

Journal ref 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08506 2021-04-20 cs.CV 57%

Visually Guided Sound Source Separation and Localization using Self-Supervised Motion Representations

Lingyu Zhu, Esa Rahtu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 18 pages. main paper: 8 pages; reference: 2 pages; supplementary material: 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.13096 2021-04-20 cs.CV 57%

Repetitive Activity Counting by Sight and Sound

Yunhua Zhang, Ling Shao, Cees G. M. Snoek

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15438 2021-03-30 cs.CV 57%

Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model

Yufan Liu, Minglang Qiao, Mai Xu, Bing Li, Weiming Hu, Ali Borji

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Published as an ECCV2020 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.05710 2021-03-30 cs.CV cs.HC 57%

Look Before you Speak: Visually Contextualized Utterances

Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.06508 2021-03-25 cs.SD cs.LG eess.AS 57%

Multi-Format Contrastive Learning of Audio Representations

Luyu Wang, Aaron van den Oord

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.11474 2021-02-24 cs.SD eess.AS 57%

Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events

Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Kai Yu

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.06837 2021-02-16 cs.CV 57%

Learning Speech-driven 3D Conversational Gestures from Video

Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, Christian Theobalt

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.05811 2021-02-12 cs.CV eess.IV 57%

Audiovisual Highlight Detection in Videos

Karel Mundnich, Alexandra Fenster, Aparna Khare, Shiva Sundaram

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 5 pages, 2 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06994 2021-02-10 eess.AS 57%

ADL-MVDR: All deep learning MVDR beamformer for target speech separation

Zhuohuang Zhang, Yong Xu, Meng Yu, Shi-Xiong Zhang, Lianwu Chen, Dong Yu

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted to ICASSP 2021, 5 pages, 2 figures; Demos are available at https://zzhang68.github.io/adlmvdr/

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.14271 2021-01-12 cs.CL 57%

Towards Fully Automated Manga Translation

Ryota Hinami, Shonosuke Ishiwatari, Kazuhiko Yasuda, Yusuke Matsui

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted to AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.01872 2021-01-07 cs.CV cs.CR 57%

Multi-Stage Residual Hiding for Image-into-Audio Steganography

Wenxue Cui, Shaohui Liu, Feng Jiang, Yongliang Liu, Debin Zhao

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14399 2021-01-07 eess.AS cs.SD 57%

Transfer Learning from Speech Synthesis to Voice Conversion with Non-Parallel Training Data

Mingyang Zhang, Yi Zhou, Li Zhao, Haizhou Li

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.01133 2020-12-03 cs.CL cs.SI 57%

It's a Thin Line Between Love and Hate: Using the Echo in Modeling Dynamics of Racist Online Communities

Eyal Arviv, Simo Hanouna, Oren Tsur

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00250 2020-12-02 cs.SD cs.HC eess.AS 57%

Strike on Stage: a percussion and media performance

Charles Martin, Chi-Hsia Lai

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Journal ref Proceedings of the International Conference on New Interfaces for Musical Expression (2011) pp. 142-143

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10283 2020-11-30 eess.AS cs.LG cs.RO cs.SD cs.SY eess.SY 57%

End-to-End Learning of Speech 2D Feature-Trajectory for Prosthetic Hands

Mohsen Jafarzadeh, Yonas Tadesse

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref 2020 Second International Conference on Transdisciplinary AI (TransAI), pages 25-33

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11387 2020-11-24 cs.CL 57%

STEPs-RL: Speech-Text Entanglement for Phonetically Sound Representation Learning

Prakamya Mishra

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.08435 2020-11-24 cs.CL 57%

SPEECH-COCO: 600k Visually Grounded Spoken Captions Aligned to MSCOCO Data Set

William Havard, Laurent Besacier, Olivier Rosec

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Data set available on https://zenodo.org/record/4282267. Presented at GLU (Grounded Language Understanding) Satellite Workshop of Interspeech 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.07340 2020-11-17 cs.CV 57%

Speech Prediction in Silent Videos using Variational Autoencoders

Ravindra Yadav, Ashish Sardana, Vinay P Namboodiri, Rajesh M Hegde

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14920 2020-10-29 cs.CL 57%

Bridging the Modality Gap for Speech-to-Text Translation

Yuchen Liu, Junnan Zhu, Jiajun Zhang, Chengqing Zong

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11226 2020-10-23 cs.SD cs.LG eess.AS 57%

Dynamic Layer Customization for Noise Robust Speech Emotion Recognition in Heterogeneous Condition Training

Alex Wilf, Emily Mower Provost

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02015 2020-10-06 cs.MM cs.GR 57%

Combined Hapto-Visual and Auditory Rendering of Cultural Heritage Objects

Praseedha Krishnan Aniyath, Sreeni Kamalalayam Gopalan, Priyadarshini K, Subhasis Chaudhuri

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM

Comments Accepted to ACCVw 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.07381 2020-09-22 cs.RO cs.AI 57%

Spatial Concept-Based Navigation with Human Speech Instructions via Probabilistic Inference on Bayesian Generative Model

Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi, Tetsunari Inamura

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to Advanced Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.02119 2020-09-07 cs.GR cs.CV cs.HC 57%

Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity

Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, Geehyuk Lee

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages; ACM Transactions on Graphics (SIGGRAPH Asia 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.12855 2020-09-01 cs.MM cs.CY cs.HC 57%

Personal Food Model

Ali Rostami, Vaibhav Pandey, Nitish Nag, Vesper Wang, Ramesh Jain

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM

Journal ref Proceedings of the 28th ACM International Conference on Multimedia (MM '20), October 12--16, 2020, Seattle, WA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05023 2020-08-13 cs.CV 57%

Audio- and Gaze-driven Facial Animation of Codec Avatars

Alexander Richard, Colin Lea, Shugao Ma, Juergen Gall, Fernando de la Torre, Yaser Sheikh

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04617 2020-08-12 eess.AS cs.SD 57%

Alzheimer's Dementia Detection from Audio and Text Modalities

Edward L. Campbell, Laura Docío-Fernández, Javier Jiménez Raboso, Carmen García-Mateo

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.03781 2020-08-11 cs.CV 57%

SemEval-2020 Task 8: Memotion Analysis -- The Visuo-Lingual Metaphor!

Chhavi Sharma, Deepesh Bhageria, William Scott, Srinivas PYKL, Amitava Das, Tanmoy Chakraborty, Viswanath Pulabaigari, Bjorn Gamback

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏