arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2011.00175 2023-11-22 eess.AS 79%

Multimodal Urban Sound Tagging with Spatiotemporal Context

Jisheng Bai, Jianfeng Chen, Mou Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11892 2023-11-21 cs.MM 79%

Multimodal Characterization of Emotion within Multimedia Space

Dayo Samuel Banjo, Connice Trimmingham, Niloofar Yousefi, Nitin Agarwal

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments 8 pages, Published in International Conference on Computers and Computation (COMPUTE 2022), November 03-04, 2022, San Francisco, United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10455 2023-11-21 eess.AS cs.SD 79%

Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement

Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Submmited to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2305.14933

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14933 2023-11-21 eess.AS cs.SD 79%

Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation

Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Published in InterSpeech 2023

Journal ref Proc. INTERSPEECH 2023, 844-848 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00641 2023-11-21 cs.CV 79%

RegBN: Batch Normalization of Multimodal Data with Regularization

Morteza Ghahremani, Christian Wachinger

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Conference on Neural Information Processing Systems (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06532 2023-11-14 cs.CL 79%

Added Toxicity Mitigation at Inference Time for Multimodal and Massively Multilingual Translation

Marta R. Costa-jussà, David Dale, Maha Elbayad, Bokai Yu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05190 2023-11-10 cs.CV 79%

Audio-visual Saliency for Omnidirectional Videos

Yuxin Zhu, Xilei Zhu, Huiyu Duan, Jie Li, Kaiwei Zhang, Yucheng Zhu, Li Chen, Xiongkuo Min, Guangtao Zhai

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 13 pages, 5 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00595 2023-10-31 cs.CV 79%

Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective

Yingying Fan, Yu Wu, Bo Du, Yutian Lin

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17568 2023-10-27 cs.HC cs.CL cs.RO 79%

Navigating to Success in Multi-Modal Human-Robot Collaboration: Analysis and Corpus Release

Stephanie M. Lukin, Kimberly A. Pollard, Claire Bonial, Taylor Hudson, Ron Arstein, Clare Voss, David Traum

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 7 pages, 3 figures

Journal ref Proceedings of the 2023 IEEE Robot and Human Interactive Communication Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11596 2023-10-26 cs.CL 79%

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, Christopher Klaiber, Pengwei Li, Daniel Licht, Jean Maillard, Alice Rakotoarison, Kaushik Ram Sadagopan, Guillaume Wenzek, Ethan Ye, Bapi Akula, Peng-Jen Chen, Naji El Hachem, Brian Ellis, Gabriel Mejia Gonzalez, Justin Haaheim, Prangthip Hansanti, Russ Howes, Bernie Huang, Min-Jae Hwang, Hirofumi Inaguma, Somya Jain, Elahe Kalbassi, Amanda Kallet, Ilia Kulikov, Janice Lam, Daniel Li, Xutai Ma, Ruslan Mavlyutov, Benjamin Peloquin, Mohamed Ramadan, Abinesh Ramakrishnan, Anna Sun, Kevin Tran, Tuan Tran, Igor Tufanov, Vish Vogeti, Carleigh Wood, Yilin Yang, Bokai Yu, Pierre Andrews, Can Balioglu, Marta R. Costa-jussà, Onur Celebi, Maha Elbayad, Cynthia Gao, Francisco Guzmán, Justine Kao, Ann Lee, Alexandre Mourachko, Juan Pino, Sravya Popuri, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, Paden Tomasello, Changhan Wang, Jeff Wang, Skyler Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11073 2023-10-17 cs.CV 79%

Audio-Visual Class-Incremental Learning

Weiguo Pian, Shentong Mo, Yunhui Guo, Yapeng Tian

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07478 2023-10-13 cs.AI 79%

Multimodal Graph Learning for Generative Tasks

Minji Yoon, Jing Yu Koh, Bryan Hooi, Ruslan Salakhutdinov

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11059 2023-10-10 eess.AS cs.SD 79%

Deep Complex U-Net with Conformer for Audio-Visual Speech Enhancement

Shafique Ahmed, Chia-Wei Chen, Wenze Ren, Chin-Jou Li, Ernie Chu, Jun-Cheng Chen, Amir Hussain, Hsin-Min Wang, Yu Tsao, Jen-Cheng Hou

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03724 2023-10-09 cs.CL 79%

Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer

Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL

Journal ref Proceedings of Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11081 2023-09-21 cs.CV 79%

Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal Distillation

Heeseung Yun, Joonil Na, Gunhee Kim

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Published to ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09709 2023-09-21 cs.CV 79%

CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video Segmentation

Kexin Li, Zongxin Yang, Lei Chen, Yi Yang, Jun Xiao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08408 2023-09-18 cs.SD eess.AS 79%

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech

Junjie Li, Ruijie Tao, Zexu Pan, Meng Ge, Shuai Wang, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Submitted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05091 2023-09-12 cs.HC cs.MM 79%

SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking Effectiveness

Zeyuan Huang, Qiang He, Kevin Maher, Xiaoming Deng, Yu-Kun Lai, Cuixia Ma, Sheng-feng Qin, Yong-Jin Liu, Hongan Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Main paper (11 pages, 6 figures) and Supplemental document (11 pages, 11 figures). Accepted by VIS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14274 2023-08-29 cs.MM 79%

Parameter-Efficient Transfer Learning for Audio-Visual-Language Tasks

Hongye Liu, Xianhai Xie, Yang Gao, Size Li, Zhou YU

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14052 2023-08-29 cs.CV 79%

MM-AU:Towards Multimodal Understanding of Advertisement Videos

Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli, Anfeng Xu, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11466 2023-08-24 cs.CL 79%

SONAR: Sentence-Level Multimodal and Language-Agnostic Representations

Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10306 2023-08-22 cs.CV 79%

Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual Navigation

Jinyu Chen, Wenguan Wang, Si Liu, Hongsheng Li, Yi Yang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12300 2023-08-21 cs.SD eess.AS 79%

A Multimodal Prototypical Approach for Unsupervised Sound Classification

Saksham Singh Kushwaha, Magdalena Fuentes

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted to INTERSPEECH 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04778 2023-08-10 cs.AI 79%

Multi-modal Multi-view Clustering based on Non-negative Matrix Factorization

Yasser Khalafaoui, Nistor Grozavu, Basarab Matei, Laurent-Walter Goix

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Journal ref 2022 IEEE Symposium Series on Computational Intelligence (SSCI), Dec 2022, Singapore, Singapore. pp.1386-1391

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04767 2023-08-10 cs.CV cs.AI cs.MM cs.SD eess.AS 79%

Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization

Tianyu Liu, Peng Zhang, Wei Huang, Yufei Zha, Tao You, Yanning Zhang

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04930 2023-08-10 cs.LG cs.CV cs.RO 79%

Multimodal Multi-User Surface Recognition with the Kernel Two-Sample Test

Behnam Khojasteh, Friedrich Solowjow, Sebastian Trimpe, Katherine J. Kuchenbecker

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01864 2023-08-02 cs.SD cs.LG eess.AS 79%

Unsupervised Improvement of Audio-Text Cross-Modal Representations

Zhepei Wang, Cem Subakan, Krishna Subramani, Junkai Wu, Tiago Tavares, Fabio Ayres, Paris Smaragdis

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted to WASPAA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15400 2023-07-31 cs.SD eess.AS 79%

The FlySpeech Audio-Visual Speaker Diarization System for MISP Challenge 2022

Li Zhang, Huan Zhao, Yue Li, Bowen Pang, Yannan Wang, Hongji Wang, Wei Rao, Qing Wang, Lei Xie

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00610 2023-07-28 cs.LG cs.CL cs.SI 79%

Fraunhofer SIT at CheckThat! 2023: Mixing Single-Modal Classifiers to Estimate the Check-Worthiness of Multi-Modal Tweets

Raphael Frick, Inna Vogel

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 8 pages

Journal ref CLEF 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13829 2023-07-27 cs.CL cs.SI 79%

ARC-NLP at Multimodal Hate Speech Event Detection 2023: Multimodal Methods Boosted by Ensemble Learning, Syntactical and Entity Features

Umitcan Sahin, Izzet Emre Kucukkaya, Oguzhan Ozcelik, Cagri Toraman

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Submitted to CASE at RANLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏