arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 1417 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 融合架构与评测 1417 篇

2107.02543 2025-12-23 cs.CV cs.HC cs.LG 57%

A Deep Learning-based Multimodal Depth-Aware Dynamic Hand Gesture Recognition System

基于深度学习的多模态深度感知动态手部姿态识别系统

Hasan Mahmud, Mashrur M. Morshed, Md. Kamrul Hasan

机构 * Systems \& Software Lab (SSL), Department of Computer Science

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本文提出了一种基于深度学习的多模态深度感知动态手部姿态识别系统,通过量化深度值和改进多模态融合CRNN架构,提高了识别准确性和参数效率。

Journal ref The Visual Computer, Springer Nature, January 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01728 2025-12-17 cs.CV 57%

Multimodal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds

基于2D正射影像和3D空中激光扫描点云的森林生物多样性潜力多模态分类

Simon B. Jensen, Stefan Oehmcke, Andreas Møgelmose, Meysam Madadi, Christian Igel, Sergio Escalera, Thomas B. Moeslund

机构 * Perception Laboratory, Aalborg University, Denmark Pioneer Centre for Artificial Intelligence, Denmark Department of Computer Science, Copenhagen University, Denmark Institute for Visual \& Analytic Computing, Rostock University, Germany University of Barcelona Computer Vision Center, Spain

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本研究利用2D正射影像和3D ALS点云数据,通过多模态深度学习融合方法,实现对森林生物多样性潜力的高效评估,实验结果达到82%的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15497 2025-12-16 eess.SP 57%

A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems

机械系统中空化强度识别的机器学习综述

Yu Sha, Ningtao Liu, Haofeng Liu, Junqi Tao, Zhenxing Niu, Guojun Huang, Yao Yao, Jiaqi Liang, Moxian Qian, Horst Stoecker, Domagoj Vnucec, Andreas Widl, Kai Zhou

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 eess.SP

AI总结 本文综述了机械系统中空化强度识别的机器学习发展,强调传统方法与深度学习的演变,并展望未来在多源数据处理和工业应用中的发展方向。

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19804 2025-12-11 cs.CV cs.AI 57%

LENVIZ: A High-Resolution Low-Exposure Night Vision Benchmark Dataset

LENVIZ:一个高分辨率低曝光夜视基准数据集

Manjushree Aithal, Rosaura G. VidalMata, Manikandtan Kartha, Gong Chen, Eashan Adhikarla, Lucas N. Kirsten, Zhicheng Fu, Nikhil A. Madhusudhana, Joe Nasti

机构 * Lenovo Research(联想研究) Lehigh University(莱文斯顿大学) Motorola Mobility(摩托罗拉移动)

专题命中 融合架构与评测 :multi-exposure(abstract);分类 cs.CV

AI总结 LENVIZ数据集为低光照图像增强提供高分辨率、多曝光的基准,包含24,000个真实场景,由专家摄影师精心制作,用于评估和改进相关技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08572 2025-12-10 cs.CV 57%

From Cells to Survival: Hierarchical Analysis of Cell Inter-Relations in Multiplex Microscopy for Lung Cancer Prognosis

从细胞到生存:多乘 microscopy 中细胞互相关系的分层分析用于肺癌预后

Olle Edgren Schüllerqvist, Jens Baumann, Joakim Lindblad, Love Nordling, Artur Mezheyeuski, Patrick Micke, Nataša Sladoje

机构 * Department of Information Technology, SciLifeLab, Uppsala University(信息科技系、SciLifeLab、乌普萨拉大学) PAICON GmbH(PAICON公司) Department of Immunology, Genetics and Pathology, Uppsala University(免疫学、遗传学和病理学系、乌普萨拉大学) Molecular Oncology Group, Vall d'Hebron Institute of Oncology(分子肿瘤学组、瓦伦·德·霍布伦肿瘤研究所) Vall d'Hebron Institute of Research(瓦伦·德·霍布伦研究所)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 HiGINE通过多光谱免疫荧光图像分析肿瘤微环境,利用分层图方法预测肺癌患者生存并提升风险分层。

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06099 2025-12-09 eess.SP 57%

Why Nonlinear Models Matter: Unified Analysis of Cognitive Load, Stress, and Exercise Using Wearable Physiological Signals

非线性模型的重要性:利用可穿戴生理信号统一分析认知负荷、压力和运动

Khondakar Ashik Shahriar

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 eess.SP

AI总结 本研究通过非线性模型统一分析认知负荷、压力和运动,证明生理状态识别的本质非线性,并建立了统一的基准以指导更稳健的可穿戴健康监测系统发展。

Comments 28 pages, 8 tables, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05593 2025-12-08 cs.CV 57%

Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer

通过无骨骼绑定图像传输学习高保真布料动画

Rong Wang, Wei Mao, Changsheng Lu, Hongdong Li

机构 * The Australian National University(澳大利亚国立大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本研究提出无骨骼绑定图像传输方法,通过独立估计顶点位置和法线以生成高保真布料动画,提升动画质量和细节恢复能力。

Comments Accepted to 3DV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19907 2025-12-08 cs.CV 57%

MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition

MHB: 多模态手形感知边界检测用于连续手语识别

Mingyu Zhao, Zhanfu Yang, Yang Zhou, Zhaoyang Xia, Can Jin, Xiaoxiao He, Dimitris N. Metaxas

机构 * Rutgers University(罗杰斯大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本文提出MHB方法,通过多模态融合结合3D骨骼特征和手形信息,提升连续手语识别的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04943 2025-12-05 cs.CV 57%

Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition

面向人类动作识别的多模态深度网络自适应融合

Novanto Yudistira

机构 * Departemen Teknik Informatika, Fakultas Ilmu Komputer, Universitas Brawijaya(计算机科学系,信息学院,布拉格亚大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本文提出了一种基于多模态深度网络的自适应融合方法,通过门控机制提升人类动作识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01750 2025-12-02 eess.SP cs.LG 57%

Multimodal Mixture-of-Experts for ISAC in Low-Altitude Wireless Networks

多模态专家混合模型用于低空无线网络中的ISAC

Kai Zhang, Wentao Yu, Hengtao He, Shenghui Song, Jun Zhang, Khaled B. Letaief

机构 * IEEE

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 eess.SP

AI总结 本文提出了一种多模态专家混合模型,用于提升低空无线网络中ISAC的性能,通过自适应融合策略提高环境感知和通信效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01519 2025-12-02 cs.CV cond-mat.mtrl-sci quant-ph 57%

QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions

QuantumCanvas:一种用于原子相互作用视觉学习的多模态基准

Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban

机构 * Texas A&M University(德克萨斯大学) Ankara University(安卡拉大学) Texas A&M University at Qatar(德克萨斯大学(卡塔尔)) Hamad Bin Khalifa University(哈马德·本·卡西姆大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 QuantumCanvas通过多模态基准统一轨道物理与视觉表示学习,为学习可转移的量子相互作用提供了系统且可解释的基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22404 2025-12-01 cs.CV 57%

UAV-MM3D: A Large-Scale Synthetic Benchmark for 3D Perception of Unmanned Aerial Vehicles with Multi-Modal Data

UAV-MM3D: 一种大规模合成基准,用于多模态数据下的无人机三维感知

Longkun Zou, Jiale Wang, Rongqin Liang, Hai Wu, Ke Chen, Yaowei Wang

机构 * Pengcheng Laboratory(鹏城实验室) University of Southern California(南加州大学)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 UAV-MM3D通过多模态合成数据提升无人机三维感知能力,提供高保真数据集和多任务基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21937 2025-12-01 cs.CV 57%

Interpretable Multimodal Cancer Prototyping with Whole Slide Images and Incompletely Paired Genomics

可解释的多模态癌症原型生成:结合整张滑片图像与不完全配对的基因组学

Yupei Zhang, Yating Huang, Wanming Hu, Lequan Yu, Hujun Yin, Chao Li

机构 * Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) Department of Electrical & Electronic Engineering, The University of Manchester, UK(曼彻斯特大学电气与电子工程系) Department of Pathology, State Key Laboratory of Oncology in South China, Guangdong Provincial Clinical Research Center for Cancer, Sun Yat-sen University Cancer Center, China(南方医科大学肿瘤学国家重点实验室、广东省癌症临床研究中心、中山大学肿瘤中心病理学部) Department of Statistics and Actuarial Science, The University of Hong Kong, Hong Kong SAR, China(香港大学统计与精算科学系) Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学学院) School of Science and Engineering and School of Medicine, University of Dundee, UK

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本文提出了一种可解释的多模态原型生成框架,通过整合整张滑片图像和不完整的基因组学数据,提升精准肿瘤学中的多模态整合效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10133 2025-11-27 cs.CV 57%

MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning

MANGO: 多模态基于注意力的归一化流方法用于融合学习

Thanh-Dat Truong, Christophe Bobda, Nitin Agarwal, Khoa Luu

机构 * CVIU Lab, University of Arkansas, USA(大学实验室,亚利桑那大学,美国) University of Florida, USA(佛罗里达大学,美国) COSMOS Research Center, University of Arkansas, Little Rock, USA(研究机构,亚利桑那大学,小石城,美国) ICSI, University of California, Berkeley, USA(研究机构,加州大学伯克利分校,美国)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 MANGO提出了一种基于归一化流和可逆交叉注意力机制的多模态融合学习方法,通过三种新型交叉注意力机制提升多模态数据的建模能力。

Comments Accepted to NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18865 2025-11-25 cs.CV 57%

DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection

DualGazeNet: 一种生物启发的双目注视查询网络用于显著物体检测

Yu Zhang, Haoan Ping, Yuchen Li, Zhenshan Bing, Fuchun Sun, Alois Knoll

机构 * School of Computation, Information and Technology, Technical University of Munich(计算信息与技术学院,慕尼黑技术大学) Department of Informatics, Technical University of Munich(信息学院,慕尼黑技术大学) University Research and Innovation Center, Obuda University(奥布达大学研究与创新中心) State Key Laboratory for Novel Software Technology, Nanjing University Suzhou Campus(新型软件技术国家重点实验室,南京大学苏州校区) School of Science and Technology, Nanjing University Suzhou Campus(科学与技术学院,南京大学苏州校区) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV

AI总结 DualGazeNet通过生物启发的纯Transformer框架,实现了显著物体检测的高效准确性和计算效率,超越了多种现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18028 2025-11-25 cs.CV 57%

MambaX: Image Super-Resolution with State Predictive Control

MambaX:基于状态预测控制的图像超分辨率

Chenyu Li, Danfeng Hong, Bing Zhang, Zhaojie Pan, Naoto Yokoya, Jocelyn Chanussot

机构 * Southeast University(东南大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) College of Resources and Environment, University of Chinese Academy of Sciences(中国科学院大学资源与环境学院) School of Mathematics, Southeast University(东南大学数学学院) Department of Complexity Science and Engineering, Graduate School of Frontier Sciences, the University of Tokyo(东京大学前沿科学研究院复杂科学与工程部门) Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、Inria、CNRS、Grenoble INP、LJK)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 MambaX通过动态状态预测控制和多模态融合方法提升图像超分辨率性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17576 2025-11-25 cs.CV cs.AI cs.LG 57%

Multimodal AI for Body Fat Estimation: Computer Vision and Anthropometry with DEXA Benchmarks

多模态AI用于脂肪率估计:计算机视觉与体态测量结合DEXA基准测试

Rayan Aldajani

机构 * Dept. of Electrical Engineering and Computer Science(电气工程与计算机科学系) York University(约克大学) Toronto, Canada(加拿大多伦多)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

AI总结 本研究提出利用多模态AI结合计算机视觉和体态测量数据,通过DEXA基准测试验证了低成本脂肪率估计方法的有效性。

Comments 2 pages, 2 figures, accepted at IEEE CASCON 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22946 2025-11-21 cs.CV 57%

LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation

LightFusion: 一种轻量级、双融合框架用于统一多模态理解和生成

Zeyu Wang, Zilong Chen, Chenhui Gou, Feng Li, Chaorui Deng, Deyao Zhu, Kunchang Li, Weihao Yu, Haoqin Tu, Haoqi Fan, Cihang Xie

机构 * UC Santa Cruz(加州大学圣克ruz分校) Tsinghua University(清华大学) Monash University(墨尔本大学) ByteDance Seed Project(字节跳动种子项目)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV

AI总结 LightFusion通过双融合机制实现轻量级多模态理解和生成,仅用350亿个标记训练即在多个基准测试中取得优异成绩。

Comments Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04638 2025-11-20 cs.CV 57%

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

Xixi Wan, Aihua Zheng, Bo Jiang, Beibei Wang, Chenglong Li, Jin Tang

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province, School of Artificial Intelligence, Anhui University(安徽省信息材料与智能感知实验室,人工智能学院,安徽大学) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University(安徽省多模态认知计算重点实验室,计算机科学与技术学院,安徽大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25238 2025-11-14 cs.CV 57%

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, Xin Jin

机构 * Nanjing University(南京大学) Huazhong University of Science and Technology(华中科技大学) Beijing Film Academy(北京电影学院) University of Science and Technology of China(中国科学技术大学) Beijing Electronic Science and Technology Institute(北京电子科技学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09773 2025-11-14 cs.LG eess.SP 57%

NeuroLingua: A Language-Inspired Hierarchical Framework for Multimodal Sleep Stage Classification Using EEG and EOG

Mahdi Samaee, Mehran Yazdi, Daniel Massicotte

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 eess.SP

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23919 2025-11-11 cs.CV 57%

Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models

Longtao Jiang, Jie Huang, Mingfei Han, Lei Chen, Yongqiang Yu, Feng Zhao, Xiaojun Chang, Zhihui Li

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05044 2025-11-10 cs.CV 57%

Medical Referring Image Segmentation via Next-Token Mask Prediction

Xinyu Chen, Yiran Wang, Gaoyang Pang, Jiafu Hao, Chentao Yue, Luping Zhou, Yonghui Li

机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院)

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE Transactions on Medical Imaging for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04347 2025-11-07 cs.CV 57%

Evaluating the Impact of Weather-Induced Sensor Occlusion on BEVFusion for 3D Object Detection

Sanjay Kumar, Tim Brophy, Eoin Martino Grua, Ganesh Sistu, Valentina Donzella, Ciaran Eising

机构 * Dept. of Electronic and Computer Engineering and the Data Driven Computer Engineering Research Centre, University of Limerick(电子与计算机工程系和数据驱动计算机工程研究中心,利默里克大学) Valeo Vision Systems(瓦莱欧视觉系统) Queen Mary University of London(伦敦女王大学)

专题命中 融合架构与评测 :sensor fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17168 2025-11-07 cs.CV 57%

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

Zhen Fan, Peng Dai, Zhuo Su, Xu Gao, Zheng Lv, Jiarui Zhang, Tianyuan Du, Guidong Wang, Yang Zhang

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18174 2025-11-05 eess.SP cs.AI cs.LG 57%

NMCSE: Noise-Robust Multi-Modal Coupling Signal Estimation Method via Optimal Transport for Cardiovascular Disease Detection

Peihong Zhang, Zhixin Li, Rui Sang, Yuxuan Liu, Yiqiang Cai, Yizhou Tan, Shengchen Li

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 eess.SP

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26681 2025-10-31 cs.CV 57%

Improving Classification of Occluded Objects through Scene Context

Courtney M. King, Daniel D. Leeds, Damian Lyons, George Kalaitzis

机构 * Fordham University Department of Computer Science(福特汉姆大学计算机科学系)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03674 2025-10-27 cs.CV cs.AI 57%

Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression

Mengshi Qi, Hao Ye, Jiaxuan Peng, Huadong Ma

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18552 2025-10-24 cs.CV 57%

Occluded nuScenes: A Multi-Sensor Dataset for Evaluating Perception Robustness in Automated Driving

Sanjay Kumar, Tim Brophy, Reenu Mohandas, Eoin Martino Grua, Ganesh Sistu, Valentina Donzella, Ciaran Eising

机构 * University of Limerick(利默里克大学) Data Driven Computer Engineering Research Centre(数据驱动计算机工程研究中心) Lero, The Irish Software Research Centre(Lero爱尔兰软件研究中心) Valeo Vision Systems(Valéo视觉系统) Queen Mary University of London(伦敦女王学院)

专题命中 融合架构与评测 :sensor fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19078 2025-10-23 cs.CV 57%

UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning

Zhongyu Jiang, Wenhao Chai, Lei Li, Zhuoran Zhou, Cheng-Yen Yang, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏