arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2406.15177 2024-06-24 cs.MM 79%

EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments ACL 2024 Demonstration Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14576 2024-06-24 eess.AS 79%

Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

Kubilay Can Demir, Belen Lojo Rodriguez, Tobias Weise, Andreas Maier, Seung Hee Yang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 5 Pages, Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12252 2024-06-19 cs.CL 79%

Language and Multimodal Models in Sports: A Survey of Datasets and Applications

Haotian Xia, Zhengbang Yang, Yun Zhao, Yuqing Wang, Jingxi Li, Rhys Tracy, Zhuangdi Zhu, Yuan-fang Wang, Hanjie Chen, Weining Shen

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10152 2024-06-17 cs.SD eess.AS 79%

Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition

Guinan Li, Jiajun Deng, Youjun Chen, Mengzhe Geng, Shujie Hu, Zhe Li, Zengrui Jin, Tianzi Wang, Xurong Xie, Helen Meng, Xunying Liu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by Interspeech 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09556 2024-06-13 eess.SP cs.AI cs.IT math.IT 79%

Co-learning-aided Multi-modal-deep-learning Framework of Passive DOA Estimators for a Heterogeneous Hybrid Massive MIMO Receiver

Jiatong Bai, Feng Shu, Qinghe Zheng, Bo Xu, Baihua Shi, Yiwen Chen, Weibin Zhang, Xianpeng Wang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06865 2024-06-12 cs.AI 79%

Eyeballing Combinatorial Problems: A Case Study of Using Multimodal Large Language Models to Solve Traveling Salesman Problems

Mohammed Elhenawy, Ahmed Abdelhay, Taqwa I. Alhadidi, Huthaifa I Ashqar, Shadi Jaradat, Ahmed Jaber, Sebastien Glaser, Andry Rakotonirainy

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06287 2024-06-10 cs.CV 79%

Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition

Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07160 2024-06-04 cs.SD cs.LG eess.AS 79%

LLark: A Multimodal Instruction-Following Language Model for Music

Josh Gardner, Simon Durand, Daniel Stoller, Rachel M. Bittner

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments ICML camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16701 2024-05-28 cs.CV 79%

Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition

Tong Shi, Xuri Ge, Joemon M. Jose, Nicolas Pugeault, Paul Henderson

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Submitted to 27th International Conference of Pattern Recognition (ICPR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13860 2024-05-24 cs.CV 79%

MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling

Diwei Huang, Kunyang Lin, Peihao Chen, Qing Du, Mingkui Tan

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 17 pages, 12 pages for main paper, 5 pages for supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04327 2024-05-08 cs.CV 79%

Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation

Dogucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Seymanur Aktı, Hazım Kemal Ekenel, Alexander Waibel

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments CVPR2024 NTIRE Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03901 2024-05-08 cs.HC cs.AI 79%

OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs

Jiahao Nick Li, Yan Xu, Tovi Grossman, Stephanie Santosa, Michelle Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Paper accepted to the 2024 CHI Conference on Human Factors in Computing Systems (CHI 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03254 2024-05-08 eess.AS 79%

Automatic Assessment of Dysarthria Using Audio-visual Vowel Graph Attention Network

Xiaokang Liu, Xiaoxia Du, Juan Liu, Rongfeng Su, Manwa Lawrence Ng, Yumei Zhang, Yudong Yang, Shaofeng Zhao, Lan Wang, Nan Yan

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 10 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03152 2024-05-08 eess.AS cs.SD 79%

MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition

Bingshen Mu, Yangze Li, Qijie Shao, Kun Wei, Xucheng Wan, Naijun Zheng, Huan Zhou, Lei Xie

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09649 2024-05-03 cs.HC cs.CL 79%

ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models

Jackie Junrui Yang, Yingtian Shi, Yuhan Zhang, Karina Li, Daniel Wan Rosli, Anisha Jain, Shuning Zhang, Tianshi Li, James A. Landay, Monica S. Lam

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17391 2024-04-29 cs.LG cs.AI cs.CY cs.HC 79%

M3BAT: Unsupervised Domain Adaptation for Multimodal Mobile Sensing with Multi-Branch Adversarial Training

Lakmal Meegahapola, Hamza Hassoune, Daniel Gatica-Perez

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at the Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (IMWUT). Paper will be presented at ACM UbiComp 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00834 2024-04-25 cs.SD cs.CV 79%

AV-RIR: Audio-Visual Room Impulse Response Estimation

Anton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya, Dinesh Manocha

专题命中 音频语音多模态 :audio-visual(title);multi-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09509 2024-04-16 cs.CV 79%

Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder

Chong Peng, Liqiang He, Dan Su

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08156 2024-04-15 cs.CL 79%

Multimodal Contextual Dialogue Breakdown Detection for Conversational AI Models

Md Messal Monem Miah, Ulie Schnaithmann, Arushi Raghuvanshi, Youngseo Son

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Published in NAACL 2024 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05559 2024-04-10 cs.CV 79%

TIM: A Time Interval Machine for Audio-Visual Action Recognition

Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman, Dima Damen

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2024. Project Webpage: https://jacobchalk.github.io/TIM-Project

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19080 2024-04-03 cs.CV cs.CR 79%

MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models

Yanting Wang, Hongye Fu, Wei Zou, Jinyuan Jia

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments To appear in CVPR'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04798 2024-04-03 cs.CL cs.LG 79%

JMI at SemEval 2024 Task 3: Two-step approach for multimodal ECAC using in-context learning with GPT and instruction-tuned Llama models

Arefa, Mohammed Abbas Ansari, Chandni Saxena, Tanvir Ahmad

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Paper Accepted at SemEval 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12687 2024-04-01 cs.CV cs.LG 79%

Audio-Visual Compound Expression Recognition Method based on Late Modality Fusion and Rule-based Decision

Elena Ryumina, Maxim Markitantov, Dmitry Ryumin, Heysem Kaya, Alexey Karpov

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19554 2024-03-29 cs.CV 79%

Cross-Attention is Not Always Needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion Recognition

R. Gnana Praveen, Jahangir Alam

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted at IEEE ICME2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18323 2024-03-28 cs.NI cs.MM 79%

How to Cache Important Contents for Multi-modal Service in Dynamic Networks: A DRL-based Caching Scheme

Zhe Zhang, Marc St-Hilaire, Xin Wei, Haiwei Dong, Abdulmotaleb El Saddik

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.MM

Journal ref IEEE Transactions on Multimedia (Early Access), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17936 2024-03-27 cs.CV 79%

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli, Marc Habermann, Christian Theobalt

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2024. Project Page: https://vcai.mpi-inf.mpg.de/projects/ConvoFusion/

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10744 2024-03-26 cs.SD cs.CV 79%

An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits

Kai Li, Fenghua Xie, Hang Chen, Kexin Yuan, Xiaolin Hu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by TPAMI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03907 2024-03-25 cs.CV 79%

Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation

Bolin Lai, Fiona Ryan, Wenqi Jia, Miao Liu, James M. Rehg

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11037 2024-03-19 eess.AS cs.SD 79%

Fine-Grained Engine Fault Sound Event Detection Using Multimodal Signals

Dennis Fedorishin, Livio Forte, Philip Schneider, Srirangaraj Setlur, Venu Govindaraju

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10689 2024-03-19 cs.RO cs.CV cs.LG 79%

Latent Object Characteristics Recognition with Visual to Haptic-Audio Cross-modal Transfer Learning

Namiko Saito, Joao Moura, Hiroki Uchida, Sethu Vijayakumar

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏