arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

1801.09056 2018-01-30 cs.CV 57%

A Multi-Biometrics for Twins Identification Based Speech and Ear

Cihan Akin, Umit Kacar, Murvet Kirci

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.07481 2018-01-24 cs.CV 57%

Survey on Emotional Body Gesture Recognition

Fatemeh Noroozi, Ciprian Adrian Corneanu, Dorota Kamińska, Tomasz Sapiński, Sergio Escalera, Gholamreza Anbarjafari

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.05197 2017-12-18 cs.IR cs.LG cs.SD eess.AS q-bio.NC 57%

Towards Deep Modeling of Music Semantics using EEG Regularizers

Francisco Raposo, David Martins de Matos, Ricardo Ribeiro, Suhua Tang, Yi Yu

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.09443 2017-09-28 cs.CL 57%

Prosodic Features from Large Corpora of Child-Directed Speech as Predictors of the Age of Acquisition of Words

Lea Frermann, Michael C. Frank

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.08168 2017-08-02 cs.CV cs.LG 57%

Look, Listen and Learn

Relja Arandjelović, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Appears in: IEEE International Conference on Computer Vision (ICCV) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.05050 2016-12-16 cs.LG cs.CV 57%

Towards Score Following in Sheet Music Images

Matthias Dorfer, Andreas Arzt, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Published In Proceedings of the 17th International Society for Music Information Retrieval Conference (2016)

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.10120 2016-12-01 cs.AI cs.HC 57%

Fusion of EEG and Musical Features in Continuous Music-emotion Recognition

Nattapong Thammasan, Ken-ichi Fukui, Masayuki Numao

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments The short version of this paper is accepted to appear as an abstract in the proceedings of AAAI-17 (student abstract and poster program)

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.01520 2016-11-22 cs.CV stat.ML 57%

An On-line Variational Bayesian Model for Multi-Person Tracking from Cluttered Scenes

Sileye Ba, Xavier Alameda-Pineda, Alessio Xompero, Radu Horaud

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 21 pages, 9 figures, 4 tables

Journal ref Computer Vision and Image Understanding, volume 153, December 2016, pages 64-76

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.02695 2016-11-10 cs.CL cs.SD 57%

Automatic recognition of child speech for robotic applications in noisy environments

Samuel Fernando, Roger K. Moore, David Cameron, Emily C. Collins, Abigail Millings, Amanda J. Sharkey, Tony J. Prescott

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Submission to Computer Speech and Language, special issue on Interaction Technologies for Children

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.08711 2016-09-01 cs.CV cs.HC 57%

Engagement Detection in Meetings

Maria Frank, Ghassem Tofighi, Haisong Gu, Renate Fruchter

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments The paper has been published on ICCCBE 2016. http://www.see.eng.osaka-u.ac.jp/seeit/icccbe2016/ http://www.see.eng.osaka-u.ac.jp/seeit/icccbe2016/download/Tentative_Time_Table_ICCCBE2016_2016-05-10.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.08955 2016-06-30 cs.MM 57%

Leveraging Contextual Cues for Generating Basketball Highlights

Vinay Bettadapura, Caroline Pantofaru, Irfan Essa

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments Proceedings of ACM Multimedia 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1408.2700 2016-04-18 cs.SD cs.MM stat.AP stat.ML 57%

Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression

Antoine Deleforge, Radu Horaud, Yoav Schechner, Laurent Girin

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments 15 pages, 8 figures

Journal ref IEEE Transactions on Audio, Speech, and Language Processing 23(4), 718-731, April, 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.2122 2016-02-22 cs.HC cs.AI cs.CY 57%

Non-Verbal Communication Analysis in Victim-Offender Mediations

Víctor Ponce-López, Sergio Escalera, Marc Pérez, Oriol Janés, Xavier Baró

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments Please, find the supplementary video material at: http://sunai.uoc.edu/~vponcel/video/VOMSessionSample.mp4

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.01042 2015-11-17 cs.CL cs.LG cs.NE 57%

Detecting Interrogative Utterances with Recurrent Neural Networks

Junyoung Chung, Jacob Devlin, Hany Hassan Awadalla

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 6 pages, accepted to NIPS 2015 Workshop on Machine Learning for Spoken Language Understanding and Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
0901.3574 2015-05-12 cs.LO cs.AI 57%

Automating Access Control Logics in Simple Type Theory with LEO-II

Christoph Benzmueller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments ii + 20 pages

Journal ref SEKI Report SR-2008-01 (ISSN 1437-4447), Saarland University, 2008

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.1411 2014-09-05 cs.CV 57%

Visual Speech Recognition

Ahmad B. A. Hassanat

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Speech and Language Technologies (Book), Prof. Ivo Ipsic (Ed.), ISBN: 978-953-307-322-4, InTech (2011)

详情

展开后加载摘要…

URL PDF HTML 收藏
1106.4451 2014-05-15 cs.MM 57%

Activities of Daily Living Indexing by Hierarchical HMM for Dementia Diagnostics

Svebor Karaman, Jenny Benois-Pineau, Jean-François Dartigues, Yann Gaëstel, Rémi Mégret, Julien Pinquier

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM

Comments 2011 9th International Workshop on Content-Based Multimedia Indexing (CBMI), Madrid : Spain (2011)

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.2124 2014-03-11 cs.CL 57%

Generating Music from Literature

Hannah Davis, Saif M. Mohammad

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL

Journal ref In Proceedings of the EACL Workshop on Computational Linguistics for Literature, April 2014, Gothenburg, Sweden

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.3475 2014-01-16 cs.LO cs.AI 57%

Prime Implicates and Prime Implicants: From Propositional to Modal Logic

Meghyn Bienvenu

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 36, pages 71-128, 2009

详情

展开后加载摘要…

URL PDF HTML 收藏
1204.4257 2012-04-20 cs.CV 57%

Speech Recognition: Increasing Efficiency of Support Vector Machines

Aamir Khan, Muhammad Farhan, Asar Ali

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 11 figures. arXiv admin note: text overlap with arXiv:1201.3720 and arXiv:1204.1177

Journal ref International Journal of Computer Applications 35(7):17-21, December 2011

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0105026 2009-11-30 cs.CV cs.HC 57%

Toward Natural Gesture/Speech Control of a Large Display

S. Kettebekov, R. Sharma

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Engineering for Human-Computer Interaction (EHCI'01),Toronto, Canada. May 11-14, 2001. Lecture Notes in Computer Science, Springer Verlag. 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0007022 2009-11-30 cs.CL 57%

ATLAS: A flexible and extensible architecture for linguistic annotation

Steven Bird, David Day, John Garofolo, John Henderson, Christophe Laprun, Mark Liberman

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments 8 pages, 9 figures

Journal ref Proceedings of the Second International Conference on Language Resources and Evaluation, pp. 1699-1706, Paris: European Language Resources Association, 2000

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08638 2025-09-16 eess.AS cs.AI cs.MM cs.SD 56%

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xinrun Du, Zhen Ye, Tianyu Zheng, Zhengxuan Jiang, Yinghao Ma, Minghao Liu, Zeyue Tian, Ziya Zhou, Liumeng Xue, Xingwei Qu, Yizhi Li, Shangda Wu, Tianhao Shen, Ziyang Ma, Jun Zhan, Chunhui Wang, Yatian Wang, Xiaowei Chi, Xinyue Zhang, Zhenzhu Yang, Xiangzhou Wang, Shansong Liu, Lingrui Mei, Peng Li, Junjie Wang, Jianwei Yu, Guojian Pang, Xu Li, Zihao Wang, Xiaohuan Zhou, Lijun Yu, Emmanouil Benetos, Yong Chen, Chenghua Lin, Xie Chen, Gus Xia, Zhaoxiang Zhang, Chao Zhang, Wenhu Chen, Xinyu Zhou, Xipeng Qiu, Roger Dannenberg, Jiaheng Liu, Jian Yang, Wenhao Huang, Wei Xue, Xu Tan, Yike Guo

机构 * HKUST(香港科技大学) MAP(多模态艺术投影)

专题命中 音频语音多模态 :分类 cs.AI、cs.MM、eess.AS;multimodal(comments)

Comments https://github.com/multimodal-art-projection/YuE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11074 2025-08-18 cs.SD cs.AI cs.CV eess.AS 56%

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters

Haomin Zhang, Kristin Qi, Shuxin Yang, Zihao Chen, Chaofan Ding, Xinhan Di

机构 * Giant Network, China(中国巨网) Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系)

专题命中 音频语音多模态 :分类 cs.CV、cs.AI、eess.AS;audio-visual(comments)

Comments Gen4AVC@ICCV: 1st Workshop on Generative AI for Audio-Visual Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.10456 2021-08-25 cs.RO cs.HC 56%

Long-Term, in-the-Wild Study of Feedback about Speech Intelligibility for K-12 Students Attending Class via a Telepresence Robot

Matthew Rueben, Mohammad Syed, Emily London, Mark Camarena, Eunsook Shin, Yulun Zhang, Timothy S. Wang, Thomas R. Groechel, Rhianna Lee, Maja J. Matarić

专题命中 音频语音多模态 :multimodal(abstract,journal_ref)

Journal ref Proceedings of the 2021 International Conference on Multimodal Interaction (ICMI '21), October 18-22, 2021, Montreal, QC, Canada. ACM, New York, NY, USA, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05700 2021-06-11 cs.HC 56%

A Wearable Virtual Touch System for Cars

Gowdham Prabhakar, Priyam Rajkhowa, Pradipta Biswas

专题命中 音频语音多模态 :multimodal(abstract,comments)

Comments Journal on Multimodal User Interface 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.13449 2020-12-29 cs.HC 56%

You Have a Point There: Object Selection Inside an Automobile Using Gaze, Head Pose and Finger Pointing

Abdul Rafey Aftab, Michael von der Beeck, Michael Feld

专题命中 音频语音多模态 :multimodal(abstract,journal_ref)

Journal ref In Proceedings of the 2020 International Conference on Multimodal Interaction, pp. 595-603. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.09197 2018-06-04 eess.AS cs.AI cs.CL cs.SD 56%

ASR-based Features for Emotion Recognition: A Transfer Learning Approach

Noé Tits, Kevin El Haddad, Thierry Dutoit

专题命中 音频语音多模态 :分类 cs.CL、cs.AI、eess.AS;multimodal(comments)

Comments Accepted to be published in the First Workshop on Computational Modeling of Human Multimodal Language - ACL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23092 2026-08-25 cs.SD 新提交 50%

Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering

面向推理的后训练与推理时LoRA重缩放:针对音频相关问答任务

Weiteng Hu, Yin Cao, Jun Yang

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 该研究针对音频相关问答任务,提出面向推理的LoRA后训练与推理时重缩放方法,在Qwen和MOSS-Audio模型上验证了有效性,提交系统在挑战赛中获总体第三、轻量级系统第二。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15206 2026-08-25 cs.LG 版本更新 50%

Chorus: Harmonizing Context and Sensing Signals for Data-Free Model Customization in IoT

Chorus:谐音上下文和传感信号以实现物联网中的无数据模型定制

Liyu Zhang, Yejia Liu, Kwun Ho Liu, Runxi Huang, Xiaomin Ouyang

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 Chorus通过学习上下文表示,在无需目标域数据的情况下,实现对未知部署条件的模型自适应,实验显示其在多种传感任务中性能优于现有方法,且推理延迟接近传感器部署。

详情

展开后加载摘要…

URL PDF HTML 收藏