arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2026-02-09 至 2026-02-09 共收录 5 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 红外-可见光融合 1 篇

2602.06363 2026-02-09 cs.CV 57%

Robust Pedestrian Detection with Uncertain Modality

具有不确定模态的鲁棒行人检测

Qian Bie, Xiao Wang, Bin Yang, Zhixi Yu, Jun Chen, Xin Xu

专题命中 红外-可见光融合 :information fusion(abstract);分类 cs.CV

AI总结 本文提出AUNet,通过自适应不确定性感知网络在不确定输入下实现鲁棒行人检测,结合UMVR和MAI模块提升跨模态信息融合效果。

Comments Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract here is shorter than that in the PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 遥感融合与全色锐化 1 篇

2602.06529 2026-02-09 cs.CV 79%

AdaptOVCD: Training-Free Open-Vocabulary Remote Sensing Change Detection via Adaptive Information Fusion

AdaptOVCD: 无需训练的开放词汇遥感变化检测 via 自适应信息融合

Mingyu Dou, Shi Qiu, Ming Hu, Yifan Chen, Huping Ye, Xiaohan Liao, Zhe Sun

机构 * Key Laboratory of Spectral Imaging Technology CAS, Xi'an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences (CAS)(光谱成像技术重点实验室、光学与精密机械研究所、中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) College of Future Information Technology, Fudan University(复旦大学未来信息技术学院) State Key Laboratory of Resources and Environment Information System, Institute of Geographic Sciences and Natural Resources Research, CAS(资源与环境信息系统国家重点实验室、地理科学与自然资源研究所、中国科学院) Key Laboratory of Low Altitude Geographic Information and Air Route, CAAC(低空地理信息与航线重点实验室、民航局) School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University(人工智能、光学与电子学院(iOPEN),西北工业大学)

专题命中 遥感融合与全色锐化 :information fusion(title,abstract);分类 cs.CV

AI总结 AdaptOVCD通过自适应信息融合实现无需训练的开放词汇遥感变化检测,有效缓解误差传播并提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 机器人多传感器融合 1 篇

2602.05538 2026-02-09 cs.CV 57%

A Comparative Study of 3D Person Detection: Sensor Modalities and Robustness in Diverse Indoor and Outdoor Environments

三维人物检测的比较研究:传感器模态与在多样室内和室外环境中的鲁棒性

Malaz Tamim, Andrea Matic-Flierl, Karsten Roscher

机构 * Fraunhofer Institute for Cognitive Systems IKS(弗劳恩霍夫认知系统研究所)

专题命中 机器人多传感器融合 :sensor fusion(abstract);分类 cs.CV

AI总结 本文比较了三种三维人物检测方法,发现融合模型在复杂场景中表现最佳,但对传感器错位仍敏感。

Comments Accepted for VISAPP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 音视频/视觉语言融合 1 篇

2602.06351 2026-02-09 cs.AI cs.CV 74%

Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion

Trifuse: 通过多模态融合增强基于注意力的GUI定位

Longhui Ma, Di Zhao, Siwei Wang, Zhao Lv, Miao Wang

机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学) Intelligent Game and Decision Lab, Academy of Military Sciences(智能游戏与决策实验室,军事科学院)

专题命中 音视频/视觉语言融合 :multimodal fusion(title);分类 cs.CV

AI总结 Trifuse通过多模态融合提升GUI定位性能,整合注意力、OCR文本和图标描述语义,无需任务微调即可实现高效定位。

Comments 17 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 融合架构与评测 1 篇

2507.06363 2026-02-09 eess.IV cs.CV 62%

Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation

Mamba Goes HoME: 分层软专家混合模型用于3D医学图像分割

Szymon Płotka, Gizem Mert, Maciej Chrabaszcz, Ewa Szczurek, Arkadiusz Sitek

机构 * Faculty of Mathematics, Informatics, and Mechanics, University of Warsaw(华沙大学数学、信息学与力学系) Faculty of Mathematics and Computer Science, Jagiellonian University(雅盖隆大学数学与计算机科学系) Institute of AI for Health, Helmholtz Munich(海德堡医学院人工智能与健康研究所) Faculty of Electronics and Information Technology, Warsaw University of Technology(华沙理工大学电子与信息技术系) NASK - National Research Institute(国家研究 institute) Faculty of Radiology, Massachusetts General Hospital(麻省总医院放射学系) Department of Radiology, Harvard Medical School(哈佛医学院放射学系)

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV、eess.IV

AI总结 本文提出分层软专家混合模型HoME,通过两级令牌路由提升3D医学图像分割的长上下文建模能力,实现更高效的局部和全局特征提取,从而提升分割性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏