arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2025-10-30 至 2025-10-30 共收录 4 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 音视频/视觉语言融合 2 篇

2510.24827 2025-10-30 cs.CV cs.MM 62%

MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition

Haoyang Zhang, Zhou Yang, Ke Sun, Yucai Pang, Guoliang Xu

机构 * Chongqing University of Posts and Telecommunications(重庆邮电大学) Xi’an Jiaotong University(西安交通大学) University of New South Wales(新南威尔士大学)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract);分类 cs.CV、cs.MM

Comments The paper will be published in the MMAsia2025 conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25070 2025-10-30 cs.CV 57%

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract);分类 cs.CV

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 融合架构与评测 2 篇

2510.24777 2025-10-30 cs.CV cs.AI eess.IV 76%

Cross-Enhanced Multimodal Fusion of Eye-Tracking and Facial Features for Alzheimer's Disease Diagnosis

Yujie Nie, Jianzhang Ni, Yonglong Ye, Yuan-Ting Zhang, Yun Kwok Wing, Xiangqing Xu, Xin Ma, Lizhou Fan

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学) Engineering Research Center of Intelligent Unmanned System, Ministry of Education(智能无人机系统工程研究中心,教育部) Department of Psychiatry, The Chinese University of Hong Kong(心理学系,香港中文大学) Department of Electronic Engineering, The Chinese University of Hong Kong(电子工程系,香港中文大学) AICARE Lab, Guangdong Medical University(AICARE实验室,广东医科大学) Department of Neurology, Shandong University of Traditional Chinese Medicine Affiliated Hospital(神经内科,山东中医药大学附属医院)

专题命中 融合架构与评测 :multimodal fusion(title);分类 cs.CV、eess.IV

Comments 35 pages, 8 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09135 2025-10-30 cs.AI cs.CL cs.HC cs.LG 71%

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L. Holt, Fernando De la Torre

机构 * Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所) Department of Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校心理学系) Center for Perceptual Systems, The University of Texas at Austin(德克萨斯大学奥斯汀分校感知系统中心)

专题命中 融合架构与评测 :multimodal fusion(title)

Comments 22 pages, first three authors equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏