arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2025-12-09 至 2025-12-09 共收录 8 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 通用Image Fusion 1 篇

2512.07170 2025-12-09 cs.CV cs.AI 85%

Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

迈向统一的语义和可控图像融合:一种扩散变换器方法

Jiayang Li, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma, Liqiang Nie

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Electronic Information School, Wuhan University(武汉大学电子信息学院)

专题命中 通用Image Fusion :image fusion(title,abstract);multi-focus(abstract);multi-exposure(abstract);分类 cs.CV

AI总结 DiTFuse通过融合图像与自然语言指令,实现端到端、语义感知的图像融合,统一了多种融合任务并在多个基准测试中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 红外-可见光融合 1 篇

2512.07760 2025-12-09 cs.CV 79%

Modality-Aware Bias Mitigation and Invariance Learning for Unsupervised Visible-Infrared Person Re-Identification

模态感知的偏见缓解与不变性学习用于无监督的可见-红外人重识别

Menglin Wang, Xiaojin Gong, Jiachen Li, Genlin Ji

专题命中 红外-可见光融合 :visible-infrared(title,abstract);分类 cs.CV

AI总结 本文提出模态感知的偏见缓解与不变性学习方法,通过改进的Jaccard距离和分割与对比策略,在无监督可见-红外人重识别中实现更可靠的跨模态关联和判别性表示学习。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 机器人多传感器融合 2 篇

2512.06848 2025-12-09 cs.CL cs.CV 74%

AquaFusionNet: Lightweight VisionSensor Fusion Framework for Real-Time Pathogen Detection and Water Quality Anomaly Prediction on Edge Devices

AquaFusionNet:轻量级视觉传感器融合框架,用于边缘设备上的实时病原体检测和水质异常预测

Sepyan Purnama Kristanto, Lutfi Hakim, Hermansyah

机构 * Department of Informatics Engineering, Politeknik Negeri Banyuwangi(信息工程系,普特里克国家理工学院巴扬威angi分校) Balai Besar Teknik Kesehatan Lingkungan dan P2B Surabaya(环境与P2B技术研究所Surabaya)

专题命中 机器人多传感器融合 :sensor fusion(title);分类 cs.CV

AI总结 AquaFusionNet通过跨模态融合提升边缘设备上病原体检测和水质异常预测的准确率与效率。

Comments 9Pages, 3 figure, Politeknik Negeri Banyuwangi

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04631 2025-12-09 cs.CV 57%

Systematic Literature Review on Vehicular Collaborative Perception -- A Computer Vision Perspective

针对车辆协同感知的系统文献综述——计算机视觉视角

Lei Wan, Jianxin Zhao, Andreas Wiedholz, Manuel Bied, Mateus Martinez de Lucena, Abhishek Dinkar Jagtap, Andreas Festag, Antônio Augusto Fröhlich, Hannan Ejaz Keen, Alexey Vinel

机构 * XITASO GmbH(XITASO公司) Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄大学) Technische Hochschule Ingolstadt (THI)(因戈尔施塔特技术大学) CARISSMA Institute of Electric, Connected and Secure Mobility (C-ECOS)(CARISSMA电动、连接与安全移动研究所) Federal University of Santa Catarina (UFSC)(圣卡塔琳娜联邦大学)

专题命中 机器人多传感器融合 :sensor fusion(abstract);分类 cs.CV

AI总结 本文从计算机视觉角度系统综述了车辆协同感知的研究现状,分析了现有方法在解决姿态误差、通信限制等实际问题中的表现,并指出现有评估指标与协同感知目标之间的不匹配。

Comments 38 pages, 8 figures, accepted for publication in IEEE Transactions on Intelligent Transportation Systems (T-ITS)

Journal ref IEEE Transactions on Intelligent Transportation Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 融合架构与评测 4 篇

2510.22410 2025-12-09 stat.AP 82%

Multimodal Fusion and Interpretability in Human Activity Recognition: A Reproducible Framework for Sensor-Based Modeling

多模态融合与可解释性在人体活动识别中的应用:一种可复现的基于传感器建模框架

Yiyao Yang, Yasemin Gulbahar

专题命中 融合架构与评测 :multimodal fusion(title,abstract);hybrid fusion(abstract)

AI总结 本文提出了一种可复现的多模态融合框架,通过统一预处理和融合策略提升人体活动识别的准确性和可解释性。

Comments 33 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06099 2025-12-09 eess.SP 57%

Why Nonlinear Models Matter: Unified Analysis of Cognitive Load, Stress, and Exercise Using Wearable Physiological Signals

非线性模型的重要性:利用可穿戴生理信号统一分析认知负荷、压力和运动

Khondakar Ashik Shahriar

专题命中 融合架构与评测 :multimodal fusion(abstract);分类 eess.SP

AI总结 本研究通过非线性模型统一分析认知负荷、压力和运动,证明生理状态识别的本质非线性,并建立了统一的基准以指导更稳健的可穿戴健康监测系统发展。

Comments 28 pages, 8 tables, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07216 2025-12-09 cs.IR cs.LG 50%

MUSE: A Simple Yet Effective Multimodal Search-Based Framework for Lifelong User Interest Modeling

MUSE:一种简单而有效的基于多模态搜索的终身用户兴趣建模框架

Bin Wu, Feifan Yang, Zhangming Chan, Yu-Ran Gu, Jiawei Feng, Chao Yi, Xiang-Rong Sheng, Han Zhu, Jian Xu, Mang Ye, Bo Zheng

机构 * Wuhan University(武汉大学) Alibaba Group(阿里巴巴集团)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 MUSE通过简单有效的多模态搜索框架,实现超长用户行为序列建模,提升推荐系统指标性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06259 2025-12-09 cs.SD cs.AI cs.LG 50%

Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling

谁会登顶排行榜?通过适应性融合模态专家和时间参与建模的多模态音乐流行度预测

Yash Choudhary, Preeti Rao, Pushpak Bhattacharyya

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 GAMENet通过融合多模态数据和时间参与建模,提升了音乐流行度预测的准确性。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏