arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2502.01709 2025-02-05 cs.SD cs.LG eess.AS 57%

Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models

Christopher Simic, Korbinian Riedhammer, Tobias Bocklet

机构 * Technische Hochschule Nuernberg(纽伦堡应用技术大学)

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted at ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18640 2025-02-03 cs.CL cs.CY cs.SI 57%

Divergent Emotional Patterns in Disinformation on Social Media? An Analysis of Tweets and TikToks about the DANA in Valencia

Iván Arcos, Paolo Rosso, Ramón Salaverría

机构 * PRHLT Research Center, Universitat Politècnica de València(瓦伦西亚理工大学PRHLT研究中心) ValgrAI Valencian Graduate School and Research Network of Artificial Intelligence(瓦伦西亚ValgrAI人工智能研究生学院及研究网络) School of Communication, Universidad de Navarra(纳瓦拉大学传播学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Journal ref Proceedings of the 17th International Conference on Agents and Artificial Intelligence (ICAART 2025), Porto, Portugal, February 23-25, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18542 2025-01-31 cs.AI 57%

Semantic Web and Creative AI -- A Technical Report from ISWS 2023

Raia Abu Ahmad, Reham Alharbi, Roberto Barile, Martin Böckling, Francisco Bolanos, Sara Bonfitto, Oleksandra Bruns, Irene Celino, Yashrajsinh Chudasama, Martin Critelli, Claudia d'Amato, Giada D'Ippolito, Ioannis Dasoulas, Stefano De Giorgis, Vincenzo De Leo, Chiara Di Bonaventura, Marco Di Panfilo, Daniil Dobriy, John Domingue, Xuemin Duan, Michel Dumontier, Sefika Efeoglu, Ruben Eschauzier, Fakih Ginwa, Nicolas Ferranti, Arianna Graciotti, Philipp Hanisch, George Hannah, Golsa Heidari, Aidan Hogan, Hassan Hussein, Alexane Jouglar, Jan-Christoph Kalo, Manoé Kieffer, Antonis Klironomos, Inês Koch, Weronika Lajewska, Nicolas Lazzari, Mikael Lindekrans, Anna Sofia Lippolis, Majlinda Llugiqi, Eleonora Mancini, Eleonora Marzi, Laura Menotti, Daniela Milon Flores, Soulakshmee Nagowah, Kerstin Neubert, Emetis Niazmand, Ebrahim Norouzi, Beatriz Olarte Martinez, Anouk Michelle Oudshoorn, Andrea Poltronieri, Valentina Presutti, Disha Purohit, Ensiyeh Raoufi, Celian Ringwald, Johanna Rockstroh, Sebastian Rudolph, Harald Sack, Zafar Saeed, Mohammad Javad Saeedizade, Aya Sahbi, Cristian Santini, Aleksandra Simic, Dennis Sommer, Rita Sousa, Mary Ann Tan, Vidyashree Tarikere, Tabea Tietz, Liam Tirpitz, Arnaldo Tomasino, Frank van Harmelen, Joao Vissoci, Caitlin Woods, Bohui Zhang, Xinyue Zhang, Heng Zheng

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13261 2025-01-24 cs.IR cs.SD eess.AS 57%

Exploring GPT's Ability as a Judge in Music Understanding

Kun Fang, Ziyu Wang, Gus Xia, Ichiro Fujinaga

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01808 2025-01-10 cs.CV 57%

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

Huaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun, Jiahui Yang, Changqing Zou, Hujun Bao

机构 * Zhejiang Lab(之江实验室) Li Auto(理想汽车) Harbin Institute of Technology(哈尔滨工业大学) Zhejiang University(浙江大学) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01700 2025-01-06 cs.CV 57%

Aesthetic Matters in Music Perception for Image Stylization: A Emotion-driven Music-to-Visual Manipulation

Junjie Xu, Xingjiao Wu, Tanren Yao, Zihao Zhang, Jiayang Bei, Wu Wen, Liang He

机构 * East China Normal University(华东师范大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00778 2025-01-03 cs.CL cs.CY 57%

Decoding the Flow: CauseMotion for Emotional Causality Analysis in Long-form Conversations

Yuxuan Zhang, Yulong Li, Zichen Yu, Feilong Tang, Zhixiang Lu, Chong Li, Kang Dang, Jionglong Su

机构 * School of Artificial Intelligence and Advanced Computing, Xi’an Jiaotong-Liverpool University(西交利物浦大学人工智能与先进计算学院) Monash University(莫纳什大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 7pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20914 2024-12-31 cs.SD cs.IR eess.AS 57%

Language-based Audio Retrieval with Co-Attention Networks

Haoran Sun, Zimu Wang, Qiuyi Chen, Jianjun Chen, Jia Wang, Haiyang Zhang

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西交利物浦大学先进技术学院)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Accepted at UIC 2024 proceedings. Accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19259 2024-12-30 eess.AS cs.SD 57%

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

Jaemin Jung, Junseok Ahn, Chaeyoung Jung, Tan Dat Nguyen, Youngjoon Jang, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12121 2024-12-30 cs.SD eess.AS 57%

WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification

Junzuo Zhou, Jiangyan Yi, Yong Ren, Jianhua Tao, Tao Wang, Chu Yuan Zhang

机构 * Institute of Automation Chinese Academy of Sciences(中国科学院自动化研究所) Department of Automation Tsinghua University(清华大学自动化系)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10523 2024-12-17 cs.CV 57%

The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

Changan Chen, Juze Zhang, Shrinidhi K. Lakshmikanth, Yusu Fang, Ruizhi Shao, Gordon Wetzstein, Li Fei-Fei, Ehsan Adeli

机构 * Stanford University(斯坦福大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Project page: languageofmotion.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15771 2024-12-17 eess.AS cs.LG cs.SD 57%

wav2pos: Sound Source Localization using Masked Autoencoders

Axel Berg, Jens Gulin, Mark O'Connor, Chuteng Zhou, Karl Åström, Magnus Oskarsson

机构 * Lund University(隆德大学) Arm(安谋科技) Sony(索尼) Tenstorrent

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments IPIN 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00571 2024-12-11 cs.SD eess.AS 57%

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview

Yupei Li, Manuel Milling, Lucia Specia, Björn W. Schuller

机构 * Imperial College London(帝国理工学院) Technical University of Munich(慕尼黑工业大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01829 2024-12-04 cs.LG cs.CV 57%

Explainable Artificial Intelligence for Medical Applications: A Review

Qiyang Sun, Alican Akman, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00949 2024-12-03 cs.LG cs.AI cs.RO 57%

STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft

Nicholas Lenzen, Amogh Raut, Andrew Melnik

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted at CoRL 2024: Workshop on Lifelong Learning for Home Robots

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12025 2024-12-02 cs.CL 57%

Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?

Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Outstanding paper at the ACL 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11369 2024-11-27 cs.SD cs.LG eess.AS 57%

Learning Spatially-Aware Language and Audio Embeddings

Bhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso, Elena Menyaylenko, Barry-John Theobald, Jonathan Sheaffer, Miguel Sarabia

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 26 pages, 7 figures, accepted at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17305 2024-11-27 cs.CV 57%

in-Car Biometrics (iCarB) Datasets for Driver Recognition: Face, Fingerprint, and Voice

Vedrana Krivokuca Hahn, Jeremy Maceiras, Alain Komaty, Philip Abbet, Sebastien Marcel

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 8 pages, 13 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13607 2024-11-27 cs.CV 57%

VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference

Seong Jong Yoo, Snehesh Shrestha, Irina Muresanu, Cornelia Fermüller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by WACV 2025 in Round 1. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15741 2024-11-26 cs.CV cs.IR cs.LG 57%

Proceedings of the 6th International Workshop on Reading Music Systems

Jorge Calvo-Zaragoza, Alexander Pacha, Elona Shatri

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings edited by Jorge Calvo-Zaragoza, Alexander Pacha and Elona Shatri

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18068 2024-11-26 cs.CV 57%

Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs

Uttaran Bhattacharya, Aniket Bera, Dinesh Manocha

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 14 pages, 7 figures, 2 tables

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 1st Workshop on Human Motion Generation, 2024, Seattle, Washington, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11454 2024-11-19 cs.CV 57%

Relevance-guided Audio Visual Fusion for Video Saliency Prediction

Li Yu, Xuanzhe Sun, Pan Gao, Moncef Gabbouj

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05550 2024-11-19 cs.HC cs.AI 57%

MEEG and AT-DGNN: Improving EEG Emotion Recognition with Music Introducing and Graph-based Learning

Minghao Xiao, Zhengxi Zhu, Kang Xie, Bin Jiang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16824 2024-11-15 cs.CV 57%

V2A-Mark: Versatile Deep Visual-Audio Watermarking for Manipulation Localization and Copyright Protection

Xuanyu Zhang, Youmin Xu, Runyi Li, Jiwen Yu, Weiqi Li, Zhipei Xu, Jian Zhang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07834 2024-11-13 cs.CV 57%

Towards Vision Mixture of Experts for Wildlife Monitoring on the Edge

Emmanuel Azuh Mensah, Anderson Lee, Haoran Zhang, Yitong Shan, Kurtis Heimerl

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16727 2024-10-31 cs.CL 57%

Recent Advances in Hate Speech Moderation: Multimodality and the Role of Large Models

Ming Shan Hee, Shivam Sharma, Rui Cao, Palash Nandi, Preslav Nakov, Tanmoy Chakraborty, Roy Ka-Wei Lee

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at EMNLP'24 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21091 2024-10-29 cs.HC cs.AI 57%

Large Language Model-assisted Speech and Pointing Benefits Multiple 3D Object Selection in Virtual Reality

Junlong Chen, Jens Grubert, Per Ola Kristensson

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19954 2024-10-29 cs.CV 57%

Turn-by-Turn Indoor Navigation for the Visually Impaired

Santosh Srinivasaiah, Sai Kumar Nekkanti, Rohith Reddy Nedhunuri

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05567 2024-10-28 cs.CV cs.HC cs.LG 57%

Exploring Emotion Expression Recognition in Older Adults Interacting with a Virtual Coach

Cristina Palmero, Mikel deVelasco, Mohamed Amine Hmani, Aymen Mtibaa, Leila Ben Letaifa, Pau Buch-Cardona, Raquel Justo, Terry Amorese, Eduardo González-Fraile, Begoña Fernández-Ruanova, Jofre Tenorio-Laranga, Anna Torp Johansen, Micaela Rodrigues da Silva, Liva Jenny Martinussen, Maria Stylianou Korsnes, Gennaro Cordasco, Anna Esposito, Mounim A. El-Yacoubi, Dijana Petrovska-Delacrétaz, M. Inés Torres, Sergio Escalera

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11368 2024-10-28 cs.RO cs.AI 57%

Dynamic Hand Gesture-Featured Human Motor Adaptation in Tool Delivery using Voice Recognition

Haolin Fei, Stefano Tedeschi, Yanpei Huang, Andrew Kennedy, Ziwei Wang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏