arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

1509.01520 2016-11-22 cs.CV stat.ML 57%

An On-line Variational Bayesian Model for Multi-Person Tracking from Cluttered Scenes

Sileye Ba, Xavier Alameda-Pineda, Alessio Xompero, Radu Horaud

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 21 pages, 9 figures, 4 tables

Journal ref Computer Vision and Image Understanding, volume 153, December 2016, pages 64-76

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.02695 2016-11-10 cs.CL cs.SD 57%

Automatic recognition of child speech for robotic applications in noisy environments

Samuel Fernando, Roger K. Moore, David Cameron, Emily C. Collins, Abigail Millings, Amanda J. Sharkey, Tony J. Prescott

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Submission to Computer Speech and Language, special issue on Interaction Technologies for Children

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.08711 2016-09-01 cs.CV cs.HC 57%

Engagement Detection in Meetings

Maria Frank, Ghassem Tofighi, Haisong Gu, Renate Fruchter

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments The paper has been published on ICCCBE 2016. http://www.see.eng.osaka-u.ac.jp/seeit/icccbe2016/ http://www.see.eng.osaka-u.ac.jp/seeit/icccbe2016/download/Tentative_Time_Table_ICCCBE2016_2016-05-10.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.08955 2016-06-30 cs.MM 57%

Leveraging Contextual Cues for Generating Basketball Highlights

Vinay Bettadapura, Caroline Pantofaru, Irfan Essa

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments Proceedings of ACM Multimedia 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1408.2700 2016-04-18 cs.SD cs.MM stat.AP stat.ML 57%

Co-Localization of Audio Sources in Images Using Binaural Features and Locally-Linear Regression

Antoine Deleforge, Radu Horaud, Yoav Schechner, Laurent Girin

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments 15 pages, 8 figures

Journal ref IEEE Transactions on Audio, Speech, and Language Processing 23(4), 718-731, April, 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.2122 2016-02-22 cs.HC cs.AI cs.CY 57%

Non-Verbal Communication Analysis in Victim-Offender Mediations

Víctor Ponce-López, Sergio Escalera, Marc Pérez, Oriol Janés, Xavier Baró

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments Please, find the supplementary video material at: http://sunai.uoc.edu/~vponcel/video/VOMSessionSample.mp4

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.01042 2015-11-17 cs.CL cs.LG cs.NE 57%

Detecting Interrogative Utterances with Recurrent Neural Networks

Junyoung Chung, Jacob Devlin, Hany Hassan Awadalla

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 6 pages, accepted to NIPS 2015 Workshop on Machine Learning for Spoken Language Understanding and Interaction

详情

展开后加载摘要…

URL PDF HTML 收藏
0901.3574 2015-05-12 cs.LO cs.AI 57%

Automating Access Control Logics in Simple Type Theory with LEO-II

Christoph Benzmueller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments ii + 20 pages

Journal ref SEKI Report SR-2008-01 (ISSN 1437-4447), Saarland University, 2008

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.1411 2014-09-05 cs.CV 57%

Visual Speech Recognition

Ahmad B. A. Hassanat

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Speech and Language Technologies (Book), Prof. Ivo Ipsic (Ed.), ISBN: 978-953-307-322-4, InTech (2011)

详情

展开后加载摘要…

URL PDF HTML 收藏
1106.4451 2014-05-15 cs.MM 57%

Activities of Daily Living Indexing by Hierarchical HMM for Dementia Diagnostics

Svebor Karaman, Jenny Benois-Pineau, Jean-François Dartigues, Yann Gaëstel, Rémi Mégret, Julien Pinquier

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM

Comments 2011 9th International Workshop on Content-Based Multimedia Indexing (CBMI), Madrid : Spain (2011)

详情

展开后加载摘要…

URL PDF HTML 收藏
1403.2124 2014-03-11 cs.CL 57%

Generating Music from Literature

Hannah Davis, Saif M. Mohammad

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL

Journal ref In Proceedings of the EACL Workshop on Computational Linguistics for Literature, April 2014, Gothenburg, Sweden

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.3475 2014-01-16 cs.LO cs.AI 57%

Prime Implicates and Prime Implicants: From Propositional to Modal Logic

Meghyn Bienvenu

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 36, pages 71-128, 2009

详情

展开后加载摘要…

URL PDF HTML 收藏
1204.4257 2012-04-20 cs.CV 57%

Speech Recognition: Increasing Efficiency of Support Vector Machines

Aamir Khan, Muhammad Farhan, Asar Ali

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 11 figures. arXiv admin note: text overlap with arXiv:1201.3720 and arXiv:1204.1177

Journal ref International Journal of Computer Applications 35(7):17-21, December 2011

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0105026 2009-11-30 cs.CV cs.HC 57%

Toward Natural Gesture/Speech Control of a Large Display

S. Kettebekov, R. Sharma

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Engineering for Human-Computer Interaction (EHCI'01),Toronto, Canada. May 11-14, 2001. Lecture Notes in Computer Science, Springer Verlag. 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0007022 2009-11-30 cs.CL 57%

ATLAS: A flexible and extensible architecture for linguistic annotation

Steven Bird, David Day, John Garofolo, John Henderson, Christophe Laprun, Mark Liberman

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments 8 pages, 9 figures

Journal ref Proceedings of the Second International Conference on Language Resources and Evaluation, pp. 1699-1706, Paris: European Language Resources Association, 2000

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08638 2025-09-16 eess.AS cs.AI cs.MM cs.SD 56%

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xinrun Du, Zhen Ye, Tianyu Zheng, Zhengxuan Jiang, Yinghao Ma, Minghao Liu, Zeyue Tian, Ziya Zhou, Liumeng Xue, Xingwei Qu, Yizhi Li, Shangda Wu, Tianhao Shen, Ziyang Ma, Jun Zhan, Chunhui Wang, Yatian Wang, Xiaowei Chi, Xinyue Zhang, Zhenzhu Yang, Xiangzhou Wang, Shansong Liu, Lingrui Mei, Peng Li, Junjie Wang, Jianwei Yu, Guojian Pang, Xu Li, Zihao Wang, Xiaohuan Zhou, Lijun Yu, Emmanouil Benetos, Yong Chen, Chenghua Lin, Xie Chen, Gus Xia, Zhaoxiang Zhang, Chao Zhang, Wenhu Chen, Xinyu Zhou, Xipeng Qiu, Roger Dannenberg, Jiaheng Liu, Jian Yang, Wenhao Huang, Wei Xue, Xu Tan, Yike Guo

机构 * HKUST(香港科技大学) MAP(多模态艺术投影)

专题命中 音频语音多模态 :分类 cs.AI、cs.MM、eess.AS;multimodal(comments)

Comments https://github.com/multimodal-art-projection/YuE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11074 2025-08-18 cs.SD cs.AI cs.CV eess.AS 56%

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters

Haomin Zhang, Kristin Qi, Shuxin Yang, Zihao Chen, Chaofan Ding, Xinhan Di

机构 * Giant Network, China(中国巨网) Computer Science, University of Massachusetts Boston(马萨诸塞大学波士顿分校计算机科学系)

专题命中 音频语音多模态 :分类 cs.CV、cs.AI、eess.AS;audio-visual(comments)

Comments Gen4AVC@ICCV: 1st Workshop on Generative AI for Audio-Visual Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.10456 2021-08-25 cs.RO cs.HC 56%

Long-Term, in-the-Wild Study of Feedback about Speech Intelligibility for K-12 Students Attending Class via a Telepresence Robot

Matthew Rueben, Mohammad Syed, Emily London, Mark Camarena, Eunsook Shin, Yulun Zhang, Timothy S. Wang, Thomas R. Groechel, Rhianna Lee, Maja J. Matarić

专题命中 音频语音多模态 :multimodal(abstract,journal_ref)

Journal ref Proceedings of the 2021 International Conference on Multimodal Interaction (ICMI '21), October 18-22, 2021, Montreal, QC, Canada. ACM, New York, NY, USA, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05700 2021-06-11 cs.HC 56%

A Wearable Virtual Touch System for Cars

Gowdham Prabhakar, Priyam Rajkhowa, Pradipta Biswas

专题命中 音频语音多模态 :multimodal(abstract,comments)

Comments Journal on Multimodal User Interface 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.13449 2020-12-29 cs.HC 56%

You Have a Point There: Object Selection Inside an Automobile Using Gaze, Head Pose and Finger Pointing

Abdul Rafey Aftab, Michael von der Beeck, Michael Feld

专题命中 音频语音多模态 :multimodal(abstract,journal_ref)

Journal ref In Proceedings of the 2020 International Conference on Multimodal Interaction, pp. 595-603. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.09197 2018-06-04 eess.AS cs.AI cs.CL cs.SD 56%

ASR-based Features for Emotion Recognition: A Transfer Learning Approach

Noé Tits, Kevin El Haddad, Thierry Dutoit

专题命中 音频语音多模态 :分类 cs.CL、cs.AI、eess.AS;multimodal(comments)

Comments Accepted to be published in the First Workshop on Computational Modeling of Human Multimodal Language - ACL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23092 2026-08-25 cs.SD 新提交 50%

Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering

面向推理的后训练与推理时LoRA重缩放:针对音频相关问答任务

Weiteng Hu, Yin Cao, Jun Yang

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 该研究针对音频相关问答任务,提出面向推理的LoRA后训练与推理时重缩放方法,在Qwen和MOSS-Audio模型上验证了有效性,提交系统在挑战赛中获总体第三、轻量级系统第二。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15206 2026-08-25 cs.LG 版本更新 50%

Chorus: Harmonizing Context and Sensing Signals for Data-Free Model Customization in IoT

Chorus:谐音上下文和传感信号以实现物联网中的无数据模型定制

Liyu Zhang, Yejia Liu, Kwun Ho Liu, Runxi Huang, Xiaomin Ouyang

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 Chorus通过学习上下文表示,在无需目标域数据的情况下,实现对未知部署条件的模型自适应,实验显示其在多种传感任务中性能优于现有方法,且推理延迟接近传感器部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20880 2026-08-24 cs.HC 新提交 50%

Live Artifacts: Authoring Dynamic Media via Live Layers Encapsulating Generative Specifications

动态制品:通过封装生成式规范的实时图层创作动态媒体

Leixian Shen, Haotian Li, Hugo Romat, Fanny Chevalier, Nicolai Marquardt, Nathalie Riche

专题命中 音频语音多模态 :cross-modal(abstract)

AI总结 该研究定义了介于静态资源与交互式软件间的Live Artifacts,提出LiveCanvas创作系统,可让创作者通过视觉画布编排动态生成式制品,经评估能帮助用户从创作静态输出转向打造响应式生成式制品。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18226 2026-08-20 cs.SD 新提交 50%

FM Synthesizer Audio-Parameter Shared Embeddings

FM合成器音频参数共享嵌入

David Braun, Adam Finkelstein

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 针对合成器预设检索问题,本文设计模仿FM信号处理的图神经网络学习含信号路由的参数表征,结合SLAP多模态目标学习音频与FM合成器参数的联合嵌入,在Yamaha DX7数据集上验证了方法的有效性与泛化性。

Comments Accepted to DAFx 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15048 2026-08-18 cs.HC 新提交 50%

Beyond Overt Reactions: Analyzing Subtle User Emotional Response to Unexpected In-Vehicle System Behavior

超越显性反应:分析用户对车载系统意外行为的细微情绪响应

Huy Quyen Ngo, Suresh Kumaar Jayaraman, Brian Mok, Ken Friedl, Oliver Krause, Aaron Steinfeld, Nikolas Martelaro

专题命中 音频语音多模态 :multi-modal(abstract)

AI总结 本研究通过驾驶模拟器收集用户与全自主车辆交互时的多模态数据,分析用户对车载系统意外行为的细微情绪响应,为设计适配乘员行为的车辆提供依据。

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15037 2026-08-18 cs.SD cs.LG 新提交 50%

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

严重声学偏移下的原型修正迭代自监督流形去噪

Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi

机构 * Indian Institute of Science Education and Research(印度科学教育与研究学院) Vellore Institute of Technology(韦洛尔理工大学)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 针对严重声学噪声下音频-文本基础模型的失效问题,提出无训练无数据源的TTA框架PRISM,在UrbanSound8K数据集上取得显著性能提升,还解决了复音陷阱问题。

Comments Accepted as a full paper at ACM CIKM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02916 2026-08-17 cs.SD cs.LG 版本更新 50%

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

SALSA-V:基于视频的捷径增强式长同步音频生成模型

Amir Dellali, Luca A. Lanzendörfer, Florian Grötschla, Roger Wattenhofer

机构 * ETH Zurich(苏黎世联邦理工学院)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 SALSA-V是一种多模态视频转音频生成模型,通过掩码扩散目标和捷径损失实现长音频同步高保真生成,仅需8步采样即可快速生成,性能优于现有SOTA,可用于专业音频合成任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12615 2026-08-14 cs.SD cs.LG 新提交 50%

Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

音乐驱动:面向车载体验的上下文感知生成音频

Cosmin Dragoiu, Nooshin Nabizadeh

机构 * Mercedes-Benz Research & Development North America(梅赛德斯-奔驰北美研发中心)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 本研究提出Drive-to-Music系统,利用行车记录仪图像与车辆遥测数据,结合感知与生成组件实现低延迟实时上下文感知车载音乐生成,为个性化自适应车载音频体验奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11496 2026-08-13 cs.LO 新提交 50%

Discrete Linear Ensemble Logic

Manfred Droste, Guo-Qiang Zhang

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏