arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 498 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 音视频/视觉语言融合 498 篇

2603.23673 2026-03-26 eess.AS cs.SD 50%

Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition

Crab:多层对比监督以提升在表演和自然语音条件下的语音情感识别

Lucas H. Ueda, João G. T. Lima, Paula D. P. Costa

机构 * Dept. of Computer Engineering and Automation (DCA), Faculdade de Engenharia Elétrica e de Computação(计算机工程与自动化系(DCA)、电气与计算机工程学院) AI Lab.(人工智能实验室) Institute of Computing, Universidade Estadual de Campinas, UNICAMP(计算学院,坎皮纳斯州立大学,UNICAMP)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出Crab模型,通过多层对比监督策略提升语音情感识别在表演和自然语音条件下的性能,采用跨模态Transformer架构和多任务对比学习,有效解决类别不平衡问题。

Comments IEEE Transactions on Affective Computing submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00030 2026-03-20 cs.LG 50%

Modality Equilibrium Matters: Minor-Modality-Aware Adaptive Alternating for Cross-Modal Memory Enhancement

模态均衡至关重要:面向跨模态记忆增强的次要模态感知自适应交替方法

Xiang Shi, Rui Zhang, Jiawei Liu, Yinpeng Liu, Qikai Cheng, Wei Lu

机构 * School of Information Management, Wuhan University(武汉大学信息管理学院)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出一种基于Shapley值的自适应交替训练框架,通过优先考虑次要模态平衡融合,引入记忆模块和跨模态映射机制,提升多模态学习性能,在四个基准数据集上取得SOTA结果。

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12583 2026-03-12 eess.AS cs.SD 50%

Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion

鲁棒的音频视觉目标说话人提取与情感感知多注册融合

Zhan Jin, Bang Zeng, Peijun Yang, Jiarong Du, Wei Ju, Yao Tian, Juan Liu, Ming Li

机构 * School of Computer Science, Wuhan University, Wuhan, China(1 武汉大学计算机学院,武汉,中国) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(2 香港中文大学(深圳)人工智能学院,中国) School of Artificial Intelligence, Wuhan University, Wuhan, China(3 武汉大学人工智能学院,武汉,中国) School of Cyber Science and Engineering, Wuhan University, Wuhan, China(4 武汉大学网络科学与工程学院,武汉,中国) Digital Innovation Research Center, Duke Kunshan University, Kunshan, China(5 香港中文大学(深圳)数字创新研究中心,中国) AI Center, OPPO, Beijing, China(6 OPPO人工智能中心,北京,中国)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出一种鲁棒的音频视觉目标说话人提取方法,通过情感感知的多注册融合技术,在模态缺失情况下提升提取性能和鲁棒性。

Comments submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08459 2026-03-10 cs.LG 50%

Data-Driven Priors for Uncertainty-Aware Deterioration Risk Prediction with Multimodal Data

数据驱动先验用于多模态数据的不确定性感知退化风险预测

L. Julián Lechuga López, Tim G. J. Rudner, Farah E. Shamout

机构 * New York University Abu Dhabi(纽约大学阿布扎克校区) University of Toronto(多伦多大学) Tandon School of Engineering, New York University(纽约大学工程学院)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

AI总结 本文提出MedCertAIn框架,通过数据驱动的先验方法提升多模态临床数据中风险预测的准确性和不确定性量化能力。

Comments 24 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01382 2026-03-03 cs.SD cs.CL 50%

End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation

端到端的同时口吃语音重建与帧级适配模块及多视图知识蒸馏

Minghui Wu, Haitao Tang, Jiahuan Fan, Ruizhi Liao, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) iFlytek Co., Ltd.(科大讯飞股份有限公司)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

AI总结 本文提出端到端的同时口吃语音重建系统,通过帧级适配模块和多视图知识蒸馏模块,提升重建语音的鲁棒性和可懂度。

Comments Submitted to 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)

Journal ref 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Singapore, 2025, pp. 1092-1097

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22961 2026-03-03 eess.AS 50%

Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models

为统一多模态语音识别适应语音基础模型与大语言模型

Jing-Xuan Zhang, Genshun Wan, Jin Li, Jianqing Gao, Duo Zhao, Zhen-Hua Ling

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出UASR-LLM框架,通过大语言模型与语音基础模型结合,实现多模态语音识别的统一优化与性能提升。

Comments 10 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23300 2026-02-27 cs.CL eess.AS 50%

A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations

一种用于对话中多模态情绪识别的专家混合模型

Soumya Dutta, Smruthi Balaji, Sriram Ganapathy

机构 * LEAP Lab, Department of Electrical Engineering(LEAP实验室,电气工程系) Microsoft(微软)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

AI总结 MiSTER-E通过专家混合框架提升对话中多模态情绪识别的准确率,实现跨模态一致性与融合。

Comments Accepted to Elsevier Computer Speech and Language. 30 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09294 2026-02-11 cs.CE 50%

BrainTAP: Brain Disorder Prediction with Adaptive Distill and Selective Prior Integration

BrainTAP: 基于自适应蒸馏和选择性先验整合的脑部疾病预测

Zhenyu Lei, Aiying Zhang, Song Wang, Han Fan, Jundong Li

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 BrainTAP通过自适应蒸馏和选择性先验整合,提升脑部疾病预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04920 2026-02-06 cs.LG cs.SD 50%

CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning

CyIN:循环信息潜在空间用于连接完整与不完整多模态学习

Ronghao Lin, Qiaolin He, Sijie Mai, Ying Zeng, Aolin Xiong, Li Huang, Yap-Peng Tan, Haifeng Hu

机构 * School of Electronics and Information Technology, Sun Yat-Sen University(中山大学电子与信息学院) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院) School of Computer Science, South China Normal University(华南师范大学计算机科学学院) Desay SV Automotive Co., Ltd(德赛股份有限公司) Pazhou Laboratory(琶洲实验室)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 CyIN通过构建循环信息潜在空间,解决多模态学习中完整与不完整数据之间的性能差距,实现统一优化。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00914 2026-02-03 cs.CL cs.AI cs.CY cs.SD eess.AS 50%

A Baseline Multimodal Approach to Emotion Recognition in Conversations

一种用于对话中情感识别的基线多模态方法

Víctor Yeste, Rodrigo Rivas-Arévalo

机构 * School of Science, Engineering and Design, Universidad Europea de Valencia(科学、工程与设计学院,欧洲大学 Valencia)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 本文提出了一种基于Transformer文本分类器和自监督语音模型的多模态基线方法,用于对话中情感识别,并通过实验展示了多模态融合的优势。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16124 2025-12-30 cs.HC cs.LG 50%

ZIA: A Theoretical Framework for Zero-Input AI

ZIA:零输入AI的理论框架

Aditi De

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥基分校)

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

AI总结 ZIA提出了一种基于多模态融合的零输入AI框架,通过整合生物信号和上下文数据实现前瞻性意图预测,提升实时推理效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17282 2025-12-04 cs.AI cs.SD 50%

ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection

ERF-BA-TFD+: 一种用于音频视觉深度伪造检测的多模态模型

Xin Zhang, Jiaming Chu, Jian Zhao, Yuchu Jiang, Xu Yang, Lei Jin, Chi Zhang, Xuelong Li

专题命中 音视频/视觉语言融合 :audio-visual fusion(abstract)

AI总结 ERF-BA-TFD+通过结合增强接收场和音频视觉融合,提出了一种多模态深度伪造检测模型,在DDL-AV数据集上实现了最先进的检测性能。

Comments The paper is withdrawn after discovering a flaw in the theoretical derivation presented in Section Method. The incorrect step leads to conclusions that are not supported by the corrected derivation. We plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19509 2025-11-26 cs.LG 50%

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

TouchFormer: 一种基于变换器的鲁棒多模态材料感知框架

Kailin Lyu, Long Xiao, Jianing Zeng, Junhao Dong, Xuexin Liu, Zhuojun Zou, Haoyue Yang, Lin Shu, Jie Hao

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

AI总结 TouchFormer通过模态自适应门控和注意力机制提升多模态材料感知的鲁棒性,并在细粒度分类任务中实现性能提升。

Comments 9 pages, 7 figures, Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02113 2025-11-11 cs.IR 50%

Enhancing Multimodal Recommendations with Vision-Language Models and Information-Aware Fusion

Hai-Dang Kieu, Min Xu, Thanh Trung Huynh, Dung D. Le

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07841 2025-11-04 cs.NI cs.LG 50%

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

Junhe Zhang, Wanli Ni, Pengwei Wang, Dongyu Wang

专题命中 音视频/视觉语言融合 :information fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27091 2025-11-03 cs.LG cs.AI quant-ph 50%

QiNN-QJ: A Quantum-inspired Neural Network with Quantum Jump for Multimodal Sentiment Analysis

Yiwei Chen, Kehuan Yan, Yu Pan, Daoyi Dong

机构 * School of Engineering, Yunnan University(云南大学工程学院) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院) Institute of Cyber-Systems and Control, College of Control Science and Engineering, Zhejiang University(浙江大学控制科学与工程学院智能系统与控制研究所) Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney(新南威尔士大学人工智能研究所)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23640 2025-10-29 cs.LG cs.AI 50%

Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning

Zihao Jing, Yan Sun, Yan Yi Li, Sugitha Janarthanan, Alana Deng, Pingzhao Hu

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23273 2025-10-28 cs.LG cs.AI q-bio.QM 50%

A Novel Framework for Multi-Modal Protein Representation Learning

Runjie Zheng, Zhen Wang, Anjie Qiao, Jiancong Xie, Jiahua Rao, Yuedong Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University (SYSU)(计算机科学与工程学院,中山大学)

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Comments 35 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17289 2025-10-21 cs.CL 50%

Addressing Antisocial Behavior in Multi-Party Dialogs Through Multimodal Representation Learning

Hajar Bakarou, Mohamed Sinane El Messoussi, Anaïs Ollagnier

机构 * Universit\'e C \ te d'Azur, CNRS, Inria, I3S Sophia Antipolis France Universit\'e C \ te d'Azur, CNRS, Inria, I3S

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08802 2025-10-13 cs.LG 50%

Edu-EmotionNet: Cross-Modality Attention Alignment with Temporal Feedback Loops

S M Rafiuddin

机构 * Department of Computer Science Oklahoma State University Stillwater, Oklahoma, USA(计算机科学系 奥克拉荷马州立大学 斯蒂尔沃特 奥克拉荷马州 美国)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments 6 Pages, 6 Figures, 3 Tables, Accepted as a Regular Research paper at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07282 2025-09-26 eess.AS 50%

Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild

Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee

专题命中 音视频/视觉语言融合 :multi-modal fusion(abstract)

Comments Proceedings of Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18706 2025-09-24 cs.HC 50%

M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition

Jiajun He, Xiaohan Shi, Cheng-Hung Hu, Jinyi Mi, Xingfeng Li, Tomoki Toda

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

Comments Accepted by IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15667 2025-09-22 cs.CL cs.SD eess.AS 50%

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros

机构 * Institute for Speech and Language Processing, Athena Research Center, Greece(语音与语言处理研究所,亚特兰蒂斯研究中心,希腊)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10307 2025-09-16 cs.IR 50%

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

Junchen Fu, Yongxin Ni, Joemon M. Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, Xuri Ge

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13072 2025-08-19 cs.AI 50%

A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis

Yuting Zhang, Tiantian Geng, Luoying Hao, Xinxing Cheng, Alexander Thorley, Xiaoxia Wang, Wenqi Lu, Sandeep S Hothi, Lei Wei, Zhaowen Qiu, Dipak Kotecha, Jinming Duan

机构 * School of Computer Science, University of Birmingham, Birmingham, UK Department of Cardiovascular Sciences, University of Birmingham, Birmingham, UK NIHR Birmingham Biomedical Research Centre West Midlands NHS Secure Data Environment, University Hospitals Birmingham NHS Foundation Trust, Birmingham, UK Department of Computing Mathematics, Manchester Metropolitan University, Manchester, UK Department of Cardiology, Heart Lung Centre, Royal Wolverhampton NHS Trust, Wolverhampton, UK Department of Cardiovascular Surgery, The First Affiliated Hospital with Nanjing Medical University , Nanjing,China College of Computer Control Engineering, Northeast Forestry University, Harbin, China Julius Center, University Medical Center Utrecht, the Netherlands Data Sciences, University of Manchester, Manchester, UK

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02133 2025-08-19 cs.HC 50%

Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs

Yitong Zhu, Lei Han, Guanxuan Jiang, PengYuan Zhou, Yuyang Wang

专题命中 音视频/视觉语言融合 :information fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20417 2025-07-29 cs.SD cs.CR eess.AS 50%

Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection

Yassine El Kheir, Arnab Das, Enes Erdem Erdogan, Fabian Ritter-Guttierez, Tim Polzehl, Sebastian Möller

机构 * Speech and Language Technology, DFKI, Germany(DFKI语音与语言技术) Quality and Usability Lab, Technical University of Berlin, Germany(柏林技术大学可用性实验室) AI Team, Gretchen AI, Germany(Gretchen AI人工智能团队) Nanyang Technological University, Singapore(南洋理工大学)

专题命中 音视频/视觉语言融合 :hybrid fusion(abstract)

Comments ACCEPTED WASPAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06247 2025-07-10 physics.flu-dyn physics.ins-det 50%

FED-PV: A Large-Scale Synthetic Frame/Event Dataset for Particle-Based Velocimetry

Fan Wu, Xiang Feng, Aoyu Zhang, Yong Lee

专题命中 音视频/视觉语言融合 :information fusion(abstract)

Comments This work has been accepted as a conference paper at the 16th International Symposium on Particle Image Velocimetry (ISPIV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12151 2025-07-08 cs.LG cs.AI 50%

Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis

Miaosen Luo, Yuncheng Jiang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22446 2025-07-01 cs.LG cs.AI 50%

EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

机构 * Dept. of Machine Learning Moffitt Cancer Center(机器学习系莫菲特癌症中心) Dept. of Cancer Epidemiology Moffitt Cancer Center(癌症流行病学系莫菲特癌症中心) Dept. of Electrical Engineering University of South Florida(电气工程系佛罗里达州立大学)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏