arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

1903.10534 2021-12-13 cs.CV cs.LG cs.SD eess.AS 62%

Learning Embodied Semantics via Music and Dance Semiotic Correlations

Francisco Afonso Raposo, David Martins de Matos, Ricardo Ribeiro

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、eess.AS

Comments 24 pages, 1 figure, 5 tables

Journal ref Neural Computing and Applications, vol. 33, pp. 14481-14493, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11984 2021-12-09 cs.SD cs.CL cs.LG eess.AS 62%

MusCaps: Generating Captions for Music Audio

Ilaria Manco, Emmanouil Benetos, Elio Quinton, Gyorgy Fazekas

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to IJCNN 2021 for the Special Session on Representation Learning for Audio, Speech, and Music Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00216 2021-12-06 cs.CV cs.SD eess.AS 62%

PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound

Zhijian Yang, Xiaoran Fan, Volkan Isler, Hyun Soo Park

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07603 2021-12-06 cs.CV cs.CL 62%

Sub-word Level Lip Reading With Visual Attention

K R Prajwal, Triantafyllos Afouras, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.06387 2021-11-12 cs.LG cs.AI cs.CV stat.ML 62%

Learning Signal-Agnostic Manifolds of Neural Fields

Yilun Du, Katherine M. Collins, Joshua B. Tenenbaum, Vincent Sitzmann

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2021, additional results and code at https://yilundu.github.io/gem/

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00404 2021-11-02 cs.SD cs.CL eess.AS 62%

Speech Emotion Recognition Using Quaternion Convolutional Neural Networks

Aneesh Muppidi, Martin Radfar

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL、eess.AS

Comments Published in ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14908 2021-11-01 cs.HC cs.CV cs.MM 62%

E-ffective: A Visual Analytic System for Exploring the Emotion and Effectiveness of Inspirational Speeches

Kevin Maher, Zeyuan Huang, Jiancheng Song, Xiaoming Deng, Yu-Kun Lai, Cuixia Ma, Hao Wang, Yong-Jin Liu, Hongan Wang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments IEEE Transactions of Visualization and Computer Graphics (TVCG, Proc. VIS 2021), to appear

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10429 2021-10-22 cs.LG cs.CL cs.SD eess.AS 62%

Knowledge distillation from language model to acoustic model: a hierarchical multi-task learning approach

Mun-Hak Lee, Joon-Hyuk Chang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 4page + 1page for citation + 2 pages for appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06280 2021-10-14 cs.SD cs.CL cs.LG eess.AS 62%

S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations

Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Hung-Yi Lee, Shinji Watanabe, Tomoki Toda

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.CL、eess.AS

Comments Submitted to ICASSP 2022. Code available at: https://github.com/s3prl/s3prl/tree/master/s3prl/downstream/a2o-vc-vcc2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02405 2021-10-07 cs.CV cs.SD eess.AS 62%

Echo-Reconstruction: Audio-Augmented 3D Scene Reconstruction

Justin Wilson, Nicholas Rewkowski, Ming C. Lin, Henry Fuchs

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08867 2021-09-22 cs.CV cs.SD eess.AS 62%

V-SlowFast Network for Efficient Visual Sound Separation

Lingyu Zhu, Esa Rahtu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments total 21 pages: main paper 8 pages, references 3 pages, supplementary material 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04236 2021-09-10 cs.CL cs.CV 62%

Identifying Visible Actions in Lifestyle Vlogs

Oana Ignat, Laura Burdick, Jia Deng, Rada Mihalcea

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted at ACL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07640 2021-08-18 cs.CV cs.SD eess.AS eess.IV 62%

Look Who's Talking: Active Speaker Detection in the Wild

You Jin Kim, Hee-Soo Heo, Soyeon Choe, Soo-Whan Chung, Yoohwan Kwon, Bong-Jin Lee, Youngki Kwon, Joon Son Chung

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments To appear in Interspeech 2021. Data will be available from https://github.com/clovaai/lookwhostalking

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01216 2021-08-17 cs.SD cs.CV eess.AS eess.IV 62%

Spot the conversation: speaker diarisation in the wild

Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments The dataset will be available for download from http://www.robots.ox.ac.uk/~vgg/data/voxceleb/voxconverse.html . The development set will be released in July 2020, and the test set will be released in October 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14102 2021-07-23 cs.CL cs.LG cs.SD eess.AS 62%

Emotion recognition by fusing time synchronous and time asynchronous representations

Wen Wu, Chao Zhang, Philip C. Woodland

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Journal ref ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 6269-6273

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12922 2021-07-21 cs.SD cs.AI cs.LG eess.AS eess.SP 62%

One Billion Audio Sounds from GPU-enabled Modular Synthesis

Joseph Turian, Jordie Shier, George Tzanetakis, Kirk McNally, Max Henry

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.08356 2021-07-20 cs.CL cs.HC cs.LG cs.MM 62%

DeHumor: Visual Analytics for Decomposing Humor

Xingbo Wang, Yao Ming, Tongshuang Wu, Haipeng Zeng, Yong Wang, Huamin Qu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM

Comments 15 pages. A preprint version of a publication at IEEE Transactions on Visualization and Computer Graphics (TVCG), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.07360 2021-07-16 cs.MM cs.HC cs.SD eess.AS 62%

Sketching sounds: an exploratory study on sound-shape associations

Sebastian Löbbers, Mathieu Barthet, György Fazekas

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments accepted for International Computer Music Conference (ICMC) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11568 2021-07-02 cs.MM cs.CV cs.LG 62%

The Influence of Audio on Video Memorability with an Audio Gestalt Regulated Video Memorability System

Lorin Sweeney, Graham Healy, Alan F. Smeaton

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 6 pages, 3 figures, 4 tables, paper accepted in CBMI 2021 for publication and oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.15684 2021-07-01 cs.CL cs.SD eess.AS 62%

Alzheimer's Dementia Recognition Using Acoustic, Lexical, Disfluency and Speech Pause Features Robust to Noisy Inputs

Morteza Rohanian, Julian Hough, Matthew Purver

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments INTERSPEECH 2021. arXiv admin note: substantial text overlap with arXiv:2106.09668

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09814 2021-06-21 cs.MM cs.SD eess.AS 62%

PixInWav: Residual Steganography for Hiding Pixels in Audio

Margarita Geleta, Cristina Punti, Kevin McGuinness, Jordi Pons, Cristian Canton, Xavier Giro-i-Nieto

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

Comments Extended abstract presented in CVPR 2021 Women in Computer Vision Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09171 2021-06-18 cs.LG cs.CV cs.SD eess.AS 62%

LiRA: Learning Visual Speech Representations from Audio through Self-supervision

Pingchuan Ma, Rodrigo Mira, Stavros Petridis, Björn W. Schuller, Maja Pantic

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、eess.AS

Comments Accepted for publication at Interspeech 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.05752 2021-06-15 cs.CL cs.LG cs.SD eess.AS 62%

Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs

Sujeong Cha, Wangrui Hou, Hyun Jung, My Phung, Michael Picheny, Hong-Kwang Kuo, Samuel Thomas, Edmilson Morais

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to Interspeech 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.03408 2021-06-01 cs.LG cs.CL eess.AS stat.ML 62%

Learning to Detect Bipolar Disorder and Borderline Personality Disorder with Language and Speech in Non-Clinical Interviews

Bo Wang, Yue Wu, Niall Taylor, Terry Lyons, Maria Liakata, Alejo J Nevado-Holgado, Kate E A Saunders

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.02725 2021-05-25 eess.AS cs.CL cs.SD 62%

Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling

Songxiang Liu, Yuewen Cao, Disong Wang, Xixin Wu, Xunying Liu, Helen Meng

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.CL、eess.AS

Comments Accepted by IEEE/ACM Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00342 2021-05-12 cs.RO cs.AI cs.LG cs.SD eess.AS 62%

Robust Robotic Pouring using Audition and Haptics

Hongzhuo Liang, Chuangchuang Zhou, Shuang Li, Xiaojian Ma, Norman Hendrich, Timo Gerkmann, Fuchun Sun, Marcus Stoffel, Jianwei Zhang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments accepted by IROS2020

Journal ref 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.08700 2021-05-06 cs.CV eess.AS 62%

A Neural Lip-Sync Framework for Synthesizing Photorealistic Virtual News Anchors

Ruobing Zheng, Zhou Zhu, Bo Song, Changjiang Ji

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、eess.AS

Comments Accepted by ICPR2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.13262 2021-04-06 cs.CV cs.AI 62%

REXUP: I REason, I EXtract, I UPdate with Structured Compositional Reasoning for Visual Question Answering

Siwen Luo, Soyeon Caren Han, Kaiyuan Sun, Josiah Poon

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICONIP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01264 2021-04-06 cs.CL cs.AI cs.LG 62%

Attention Forcing for Machine Translation

Qingyun Dou, Yiting Lu, Potsawee Manakul, Xixin Wu, Mark J. F. Gales

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:1909.12289

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10018 2021-03-19 cs.SD cs.MM eess.AS 62%

Audio Description from Image by Modal Translation Network

Hailong Ning, Xiangtao Zheng, Yuan Yuan, Xiaoqiang Lu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏