arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-06 至 2026-02-06 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 6 篇

2602.03891 2026-02-06 eess.AS cs.AI cs.CV cs.MM cs.SD 83%

Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection

音频亮点:用于音频-视觉视频亮点检测的双路径音频编码器

Seohyun Joo, Yoori Oh

机构 * School of Electrical Engineering(电气工程学院) Music and Audio Research Group(音乐与音频研究组) Seoul National University(首尔国立大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出双路径音频编码器,通过结合语义和动态路径,提升音频-视觉视频亮点检测的性能。

Comments 5 pages, 2 figures, to appear in ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04924 2026-02-06 cs.LG cs.SD 82%

Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering

何时回答:可靠音频-视觉问答的自适应置信度细化

Dinh Phu Tran, Jihoon Jeong, Saad Wazir, Seongah Kim, Thao Do, Cem Subakan, Daeyoung Kim

机构 * School of Computing, KAIST, Republic of Korea(韩国釜山科学技术院计算机科学学院) Laval University(拉瓦尔大学) Mila-Quebec AI Institute(魁北克人工智能研究所)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract)

AI总结 本文提出自适应置信度细化方法,用于提升音频-视觉问答任务中可靠性的性能。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05093 2026-02-06 cs.HC 67%

VR Calm Plus: Coupling a Squeezable Tangible Interaction with Immersive VR for Stress Regulation

VR Calm Plus:将可压缩的实体交互与沉浸式VR结合用于压力调节

He Zhang, Xinyang Li, Xingyu Zhou, Xinyi Fu

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract)

AI总结 VR Calm Plus通过结合可压缩实体交互与沉浸式VR,提升压力调节效果,增强积极情绪与放松体验。

Comments This work has been conditionally accepted by the ACM CHI 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08176 2026-02-06 cs.SD cs.AI cs.LG eess.AS 62%

Leveraging Whisper Embeddings for Audio-based Lyrics Matching

利用Whisper嵌入进行基于音频的歌词匹配

Eleonora Mancini, Joan Serrà, Paolo Torroni, Yuki Mitsufuji

机构 * DISI, University of Bologna(博洛尼亚大学DISI) Sony AI(索尼人工智能) Sony Group Corporation(索尼集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出WEALY,利用Whisper嵌入实现基于音频的歌词匹配,通过可重复的流程和多模态扩展,展示了与最先进方法相当的性能,并为音乐信息检索提供了可靠基准。

Comments Accepted at ICASSP 2026 (IEEE International Conference on Acoustics, Speech and Signal Processing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05107 2026-02-06 cs.CL 57%

Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text

多语言语音与文本中隐含语篇关系的提取与识别

Ahmed Ruby, Christian Hardmeier, Sara Stymne

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文提出了一种多模态方法,通过整合文本和音频信息,实现多语言隐含语篇关系的跨语言迁移和识别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04913 2026-02-06 cs.LG cs.AI cs.SD 57%

A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model

A$^2$-LLM:一种端到端的对话音频虚拟形象大语言模型

Xiaolin Hu, Hang Yuan, Xinzhu Sang, Binbin Yan, Zhou Yu, Cong Huang, Kai Chen

机构 * State Key Laboratory of Information Photonics(信息光子学国家重点实验室) Optical Communications, Beijing University of Posts(邮电大学光学通信学院) Zhongguancun Institute of Artificial Intelligence, Beijing, China(中关村人工智能研究院) School of Finance and Statistics, East China Normal University, Shanghai, China(东华大学金融与统计学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 A$^2$-LLM通过统一框架联合推理语言、音频语调和3D面部运动,实现端到端的情感表达对话虚拟形象,提升实时效率与情感表现力。

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏