arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-03 至 2026-02-03 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 8 篇

2602.00701 2026-02-03 cs.MM cs.CV cs.LG cs.SD 90%

Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning

跨模态二进制注意力:面向音频视觉学习的高效融合框架

Mohamed Saleh, Zahra Ahmadi

机构 * Peter L. Reichertz Institute for Medical Informatics of TU Braunschweig and Hannover Medical School(图林根工业大学和汉诺威医学院医学信息学研究所) Lower Saxony Center for AI and Causal Methods in Medicine (CAIMed)(下萨克森人工智能与因果医学方法中心)

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

AI总结 提出CMQKA和SNNergy,通过高效二进制操作实现线性复杂度的跨模态融合,显著提升音频视觉任务的能耗效率和性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18110 2026-02-03 cs.CL 86%

Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM

Watch and Listen: 通过多模态大语言模型理解音频-视觉-语音时刻

Zinuo Li, Xian Zhang, Yongxin Guo, Mohammed Bennamoun, Farid Boussaid, Girish Dwivedi, Luqi Gong, Qiuhong Ke

机构 * University of Western Australia(西澳大学) Alibaba Group(阿里巴巴集团) Zhejiang Laboratory(浙江实验室) Monash University(墨尔本大学)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title);分类 cs.CL

AI总结 TriSense通过整合视觉、音频和语音模态,提升视频时间理解能力,采用基于查询的连接器实现多模态鲁棒性,并通过TriSense-2M数据集推动多模态视频分析发展。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01284 2026-02-03 cs.MM cs.CV cs.HC 81%

Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection

看见、听见与认知:深度伪造视频检测中的多模态策略

Chen Chen, Dion Hoe-Lian Goh

机构 * Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 研究探讨了多模态策略在深度伪造视频检测中的应用,通过分析人类识别过程中的线索组合,提出了提升媒体素养的指导方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00189 2026-02-03 cs.SD cs.AI cs.MM eess.AS 67%

LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild

LPIPS-AttnWav2Lip:通用音频驱动的唇同步生成野态说话人面部图像

Zhipeng Chen, Xinheng Wang, Lun Xie, Haijie Yuan, Hang Pan

机构 * University of Science and Technology Beijing(北京科技大学) Xiaoduo Intelligent Technology (Beijing) Co., Ltd(小多智能技术(北京)有限公司) Department of Computer Science, Changzhi University(长治大学计算机科学系)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 本文提出LPIPS-AttnWav2Lip方法,通过音频驱动生成逼真高质的说话人面部图像,实现精确唇同步和视觉质量提升。

Comments This paper has been accepted by Elsevier's \textit{Speech Communication} journal. Official publication link: https://doi.org/10.1016/j.specom.2023.103028 The code for the paper is available at the following link: https://github.com/FelixChan9527/LPIPS-AttnWav2Lip

Journal ref Speech Communication 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02741 2026-02-03 cs.LG cs.AI cs.SD eess.AS 62%

DeepGB-TB: A Risk-Balanced Cross-Attention Gradient-Boosted Convolutional Network for Rapid, Interpretable Tuberculosis Screening

DeepGB-TB: 一种风险平衡的跨注意力梯度提升卷积网络用于快速、可解释的肺结核筛查

Zhixiang Lu, Yulong Li, Feilong Tang, Zhengyong Jiang, Chong Li, Mian Zhou, Tenglong Li, Jionglong Su

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

AI总结 DeepGB-TB通过结合跨注意力机制和梯度提升决策树,实现快速、可解释的肺结核筛查,具有高准确率和低资源需求。

Comments Accepted by AAAI 2026 (oral)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02100 2026-02-03 cs.CY cs.AI cs.SI 57%

The Verification Crisis: Expert Perceptions of GenAI Disinformation and the Case for Reproducible Provenance

生成式人工智能的验证危机:专家对生成式人工智能虚假信息的看法及可重复性溯源的案例

Alexander Loth, Martin Kappes, Marc-Oliver Pahl

机构 * Frankfurt University of Applied Sciences(法兰克福应用科学大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文探讨生成式人工智能虚假信息的验证危机,指出大规模文本生成带来的系统性风险,并提出通过可重复性溯源和监管框架来应对挑战。

Comments Accepted at ACM TheWebConf '26 Companion

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01547 2026-02-03 cs.SD eess.AS 57%

Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition

基于注意力加权中心核对齐的知识蒸馏:用于大音频-语言模型的语音情感识别

Qingran Yang, Botao Zhao, Zuheng Kang, Xue Li, Yayun He, Chuhang Liu, Xulong Zhang, Xiaoyang Qu, Junqing Peng, Jianzong Wang

机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

AI总结 PL-Distill通过结合投影层和输出层蒸馏方法,有效压缩大音频-语言模型并提升语音情感识别性能。

Comments Accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07213 2026-02-03 cs.AI 57%

Brain-inspired Computing Based on Deep Learning for Human-computer Interaction: A Review

基于深度学习的脑启发计算用于人机交互:综述

Bihui Yu, Sibo Zhang, Lili Zhou, Jingxuan Wei, Linzhuang Sun, Liping Bu

机构 * Shenyang Institute of Computing Technology, Chinese Academy of Sciences(中国科学院沈阳计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) Heilongjiang Academy of Sciences Intelligent Manufacturing Institute(黑龙江科学院智能制造研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文综述了基于深度学习的脑启发计算在人机交互中的应用,探讨了其发展、挑战及未来研究方向。

Comments 26pages, 8 figures and 4 tables

Journal ref Neurocomputing, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏