arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

1811.04448 2018-11-13 cs.SD eess.AS 79%

A Multi-modal Deep Neural Network approach to Bird-song identification

Botond Fazeka, Alexander Schindler, Thomas Lidy, Andreas Rauber

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments LifeCLEF 2017 working notes, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.09725 2018-10-15 cs.CV cs.SD 79%

Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion

Israel D. Gebru, Silèye Ba, Xiaofei Li, Radu Horaud

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 14 pages, 6 figures, 5 tables

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6), 1086 - 1099, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.00108 2018-10-02 cs.CV 79%

Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture

Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Georgios Tzimiropoulos, Maja Pantic

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to IEEE SLT 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.06708 2018-09-18 cs.CV 79%

Speech Map: A Statistical Multimodal Atlas of 4D Tongue Motion During Speech from Tagged and Cine MR Images

Jonghye Woo, Fangxu Xing, Maureen Stone, Jordan Green, Timothy G. Reese, Thomas J. Brady, Van J. Wedeen, Jerry L. Prince, Georges El Fakhri

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at Journal of Computer Methods in Biomechanics and Biomedical Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.04203 2018-09-13 cs.CL 79%

Multimodal neural pronunciation modeling for spoken languages with logographic origin

Minh Nguyen, Gia H. Ngo, Nancy F. Chen

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.05561 2018-08-17 cs.CV 79%

Emotion Recognition in Speech using Cross-Modal Transfer in the Wild

Samuel Albanie, Arsha Nagrani, Andrea Vedaldi, Andrew Zisserman

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Conference paper at ACM Multimedia 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.00118 2018-08-02 cs.CV cs.CR cs.HC 79%

Toward Multimodal Interaction in Scalable Visual Digital Evidence Visualization Using Computer Vision Techniques and ISS

Serguei A. Mokhov, Miao Song, Jashanjot Singh, Joey Paquet, Mourad Debbabi, Sudhir Mudur

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments reformatted; ICPRAI 2018 conference proceedings, pp. 151-157, CENPARMI, Concordia University, Montreal

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.08612 2018-06-25 cs.CV 79%

Ad-Net: Audio-Visual Convolutional Neural Network for Advertisement Detection In Videos

Shervin Minaee, Imed Bouazizi, Prakash Kolan, Hossein Najafzadeh

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04121 2018-06-20 cs.CV cs.SD 79%

The Conversation: Deep Audio-Visual Speech Enhancement

Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments To appear in Interspeech 2018. We provide supplementary material with interactive demonstrations on http://www.robots.ox.ac.uk/~vgg/demo/theconversation

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.02923 2018-06-11 cs.CL 79%

Multimodal Relational Tensor Network for Sentiment and Emotion Classification

Saurav Sahay, Shachi H Kumar, Rui Xia, Jonathan Huang, Lama Nachman

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00326 2018-04-04 cs.CV 79%

Seeing Voices and Hearing Faces: Cross-modal biometric matching

Arsha Nagrani, Samuel Albanie, Andrew Zisserman

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments To appear in: IEEE Computer Vision and Pattern Recognition (CVPR), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.08332 2018-02-26 cs.CL 79%

Deep Multimodal Learning for Emotion Recognition in Spoken Language

Yue Gu, Shuhong Chen, Ivan Marsic

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments ICASSP 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.01115 2018-02-06 cs.CV 79%

End2You -- The Imperial Toolkit for Multimodal Profiling by End-to-End Learning

Panagiotis Tzirakis, Stefanos Zafeiriou, Bjorn W. Schuller

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.11118 2017-12-01 cs.CL 79%

Multimodal Attribute Extraction

Robert L. Logan, Samuel Humeau, Sameer Singh

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments AKBC 2017 Workshop Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.08976 2017-11-30 cs.IR cs.SD eess.AS 79%

Deep Cross-Modal Correlation Learning for Audio and Lyrics in Music Retrieval

Yi Yu, Suhua Tang, Francisco Raposo, Lei Chen

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.01775 2017-11-07 cs.MM cs.HC cs.RO 79%

Multimodal Signal Processing and Learning Aspects of Human-Robot Interaction for an Assistive Bathing Robot

A. Zlatintsi, I. Rodomagoulakis, P. Koutras, A. C. Dometios, V. Pitsikalis, C. S. Tzafestas, P. Maragos

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02757 2017-06-12 cs.RO cs.CL cs.HC 79%

Sympathy Begins with a Smile, Intelligence Begins with a Word: Use of Multimodal Features in Spoken Human-Robot Interaction

Jekaterina Novikova, Christian Dondrup, Ioannis Papaioannou, Oliver Lemon

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Robo-NLP workshop at ACL 2017. 9 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.01509 2017-01-31 cs.CV cs.LG stat.ML 79%

EM Algorithms for Weighted-Data Clustering with Application to Audio-Visual Scene Analysis

Israel D. Gebru, Xavier Alameda-Pineda, Florence Forbes, Radu Horaud

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 14 pages, 4 figures, 4 tables

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, volume 38, number 12, 2402 - 2415, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08666 2016-11-29 cs.LG cs.AI cs.RO 79%

Training an Interactive Humanoid Robot Using Multimodal Deep Reinforcement Learning

Heriberto Cuayáhuitl, Guillaume Couly, Clément Olalainty

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments NIPS Workshop on Future of Interactive Learning Machines, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.02946 2016-11-23 cs.CV 79%

Kernel-based Sensor Fusion with Application to Audio-Visual Voice Activity Detection

David Dov, Ronen Talmon, Israel Cohen

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1303.5395 2013-03-25 cs.AI 79%

Lattice-Based Graded Logic: a Multimodal Approach

Philippe Chatalic, Christine Froidevaux

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Appears in Proceedings of the Eighth Conference on Uncertainty in Artificial Intelligence (UAI1992)

详情

展开后加载摘要…

URL PDF HTML 收藏
1201.3720 2012-01-19 cs.CV 79%

A Multimodal Biometric System Using Linear Discriminant Analysis For Improved Performance

Aamir Khan, Muhammad Farhan, Aasim Khurshid, Adeel Akram

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref IJCSI International Journal of Computer Science Issues, Vol. 8, Issue 6, No 2, 2011, 122-127

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.6361 2011-09-30 cs.AI 79%

Cognitive Principles in Robust Multimodal Interpretation

J. Y. Chai, Z. Prasov, S. Qu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 27, pages 55-83, 2006

详情

展开后加载摘要…

URL PDF HTML 收藏
1005.4014 2010-05-24 cs.MM 79%

A Study on Potential of Integrating Multimodal Interaction into Musical Conducting Education

Gilbert Phuah Leong Siang, Nor Azman Ismail, Pang Yee Yong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments http://www.journalofcomputing.org

Journal ref Journal of Computing, Volume 2, Issue 5, May 2010

详情

展开后加载摘要…

URL PDF HTML 收藏
cmp-lg/9406002 2009-11-30 cmp-lg cs.CL 79%

Speech Dialogue with Facial Displays: Multimodal Human-Computer Conversation

Katashi Nagao, Akikazu Takeuchi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 8 pages, Postscript file, UNIX compressed, uuencoded

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11007 2026-07-29 cs.LG 版本更新 79%

TabPFN beyond Tabular Data: Calibration and Accuracy on Multimodal Embeddings

超越表格数据的TabPFN:多模态嵌入的校准与准确性

Jingxiang Zhang, Lujia Zhong, Zijie Zhu, Shuo Huang, Yuang Xu

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 研究少样本多模态分类中轻量级头部校准不佳问题,核心方法是评估TabPFN作为即插即用分类头。贡献在于其在多数据集、多编码器和多模态评估中表现优异,能提升校准并保持准确率,为校准敏感多模态分类提供指导。

Comments 19 pages, 13 figures, 10 tables. Jingxiang Zhang and Lujia Zhong contributed equally. Code: https://github.com/Jingxiang-Zhang/tabpfn-multimodal-embeddings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07367 2026-03-03 cs.SD 79%

FOCAL: A Novel Benchmarking Technique for Multi-modal Agents

FOCAL:多模态代理的新型基准测试技术

Anupam Purwar, Aditya Choudhary

专题命中 音频语音多模态 :multi-modal(title,abstract)

AI总结 FOCAL提出了一种新型基准测试技术,用于评估多模态代理的端到端推理能力、误差传播及对话效果。

Comments We present a framework for evaluation of Multi-modal Agents consisting of Voice-to-voice model components viz. Text to Speech (TTS), Retrieval Augmented Generation (RAG) and Speech-to-text (STT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25041 2026-06-26 cs.CV cs.AI cs.GR cs.SD 新提交 79%

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Wan-Streamer v0.1: 端到端实时交互基础模型

Lianghua Huang, Zhi-Fan Wu, Wei Wang, Yupeng Shi, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuzheng Wang, Zoubin Bi

机构 * Alibaba Group(阿里巴巴集团)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 提出Wan-Streamer,一种原生流式、端到端的交互基础模型,通过单一Transformer联合建模语言、音频和视频,实现低延迟全双工音视频交互,无需外部模块。

Comments Website: https://wan-streamer.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23552 2026-05-18 cs.CV cs.SD eess.AS 79%

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching

JAM-Flow:联合音频-运动合成与流匹配

Mingi Kwon, Joonghyuk Shin, Jaeseok Jung, Jaesik Park, Youngjung Uh

机构 * Yonsei University(延世大学) CineLingo Seoul National University(首尔国立大学)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、eess.AS

AI总结 本文提出JAM-Flow框架,通过流匹配和多模态扩散变压器架构,实现面部运动与语音的联合合成与条件生成,支持文本到说话人生成、音频驱动动画等多种任务。

Comments project page: https://joonghyuk.com/jamflow-web Under review. Preprint published on arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11605 2026-05-13 cs.CV cs.AI 79%

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

保留音频无法表达的内容:面向多模态大语言模型的上下文保留令牌修剪

Chaeyoung Jung, Kyeongha Rho, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ContextGuard框架,通过保留音频-视觉上下文并去除跨模态冗余,实现多模态大语言模型的高效令牌修剪,实验表明其在多个基准测试中表现优异,能有效减少输入令牌数量。

详情

展开后加载摘要…

URL PDF HTML 收藏