arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

2025-08-12 至 2025-08-12 共收录 11 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 通用Image Fusion 2 篇

2508.07803 2025-08-12 cs.CV 87%

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng, Haishu Tan, Huafeng Li

机构 * School of Physics and Optoelectronic Engineering(物理与光电工程学院) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) School of Information Engineering and Automation(信息工程与自动化学院)

专题命中 通用Image Fusion :multimodal fusion(title,abstract);image fusion(abstract);multimodal image fusion(abstract);infrared and visible(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04377 2025-08-12 cs.CV cs.CL cs.MM 62%

Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

Xiao Zhang, Johan Bos

机构 * University of Groningen(格罗宁根大学)

专题命中 通用Image Fusion :image fusion(abstract);分类 cs.CV、cs.MM

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 遥感融合与全色锐化 1 篇

2508.08183 2025-08-12 cs.CV eess.IV 81%

THAT: Token-wise High-frequency Augmentation Transformer for Hyperspectral Pansharpening

Hongkun Jin, Hongcheng Jiang, Zejun Zhang, Yuan Zhang, Jia Fu, Tingfeng Li, Kai Luo

机构 * JPMorgan Chase(摩根大通公司) University of Missouri-Kansas City(密苏里大学-堪萨斯城分校) Ming Hsieh Department of Electrical and Computer Engineering(明希赫电气与计算机工程系) University of Southern California(南加州大学) Robinson Research Institute(罗宾逊研究学院) KTH Royal Institute of Technology(皇家理工学院) NEC Laboratories America(NEC美国实验室) University of Virginia(弗吉尼亚大学)

专题命中 遥感融合与全色锐化 :pansharpening(title,abstract);分类 cs.CV、eess.IV

Comments Accepted to 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 自动驾驶多传感器融合 2 篇

2508.07560 2025-08-12 cs.RO cs.CV 73%

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang, Yongsheng Gao, Jie Zhao, Zifan Huang, Haozhi Bai, Nanxin Zeng, Nayu Su, Lei Yang, Ziying Song, Xiaoxi Hu, Xinmin Jiang, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System(机器人系统国家重点实验室) Harbin Institute of Technology(哈尔滨工业大学) State Key Laboratory of Intelligent Green Vehicle and Mobility(智能绿色车辆与移动性国家重点实验室) Tsinghua University(清华大学) the School of Mechanical and Aerospace Engineering(机械与航空航天工程学院) Nanyang Technological University(南洋理工大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) Beijing Jiaotong University(北京交通大学) the Institute for Infocomm Research(信息通信研究所) A*STAR the Engineering Cluster(工程集群) the Singapore Institute of Technology(新加坡理工学院)

专题命中 自动驾驶多传感器融合 :sensor fusion(abstract);multi-sensor fusion(abstract);分类 cs.CV、cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07453 2025-08-12 eess.SY cs.AI cs.MA cs.RO cs.SY 57%

Noise-Aware Generative Microscopic Traffic Simulation

Vindula Jayawardana, Catherine Tang, Junyi Ji, Jonah Philion, Xue Bin Peng, Cathy Wu

机构 * Massachusetts Institute of Technology(麻省理工学院) Vanderbilt University(范德比大学) NVIDIA Corporation(NVIDIA公司)

专题命中 自动驾驶多传感器融合 :sensor fusion(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 机器人多传感器融合 1 篇

2508.06566 2025-08-12 cs.CV cs.AI 57%

Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features

Manish Kansana, Elias Hossain, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science and Engineering, Mississippi State University(计算机科学与工程系,密苏里州立大学)

专题命中 机器人多传感器融合 :multimodal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 音视频/视觉语言融合 2 篇

2508.06902 2025-08-12 cs.CV 79%

eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos

Xuecheng Wu, Dingkang Yang, Danlei Huang, Xinyi Yin, Yifan Wang, Jia Zhang, Jiayu Nie, Liangyu Fu, Yang Liu, Junxiao Xue, Hadi Amirpour, Wei Zhou

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Intelligent Robotics and Advanced Manufacturing, Fudan University & ByteDance(智能机器人与先进制造学院,复旦大学 & 字节跳动) School of Cyber Science and Engineering, Zhengzhou University(网络科学与工程学院,郑州大学) Institute of Advanced Technology, University of Science and Technology of China(先进技术研究院,中国科学技术大学) Inspur Electronic Information Industry Co., Ltd(Inspur电子信息产业有限公司) Department of Computer Science, The University of Toronto(计算机科学系,多伦多大学) Research Center for Space Computing System, Zhejiang Lab(空间计算系统研究中心,浙江实验室) Institute of Information Technology, University of Klagenfurt(信息技术学院,克雷格福特大学) School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学)

专题命中 音视频/视觉语言融合 :audio-visual fusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07023 2025-08-12 cs.CV 57%

MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering

Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas, Meilin Zhang

机构 * Shaanxi University of Technology(陕西理工大学) Technological University of Peru(秘鲁技术大学)

专题命中 音视频/视觉语言融合 :multimodal fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 融合架构与评测 3 篇

2508.08075 2025-08-12 cs.AI 78%

FNBT: Full Negation Belief Transformation for Open-World Information Fusion Based on Dempster-Shafer Theory of Evidence

Meishen He, Wenjun Ma, Jiao Wang, Huijun Yue, Xiaoma Fan

专题命中 融合架构与评测 :information fusion(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06701 2025-08-12 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 74%

MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆斯美国大学)

专题命中 融合架构与评测 :multimodal fusion(title);分类 cs.CV

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10679 2025-08-12 cs.CV 57%

Alignment-free Raw Video Demoireing

Shuning Xu, Xina Liu, Binbin Song, Xiangyu Chen, Qiubo Chen, Jiantao Zhou

专题命中 融合架构与评测 :information fusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏