arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

多模态信息融合

面向图像、视频、多传感器和跨模态感知的信息融合,包括 Image Fusion、红外可见光、遥感、医学影像、LiDAR/雷达/相机和音视频融合。

共收录 1417 信号源:cs.CV, eess.IV, eess.SP, cs.RO, cs.MM

1. 融合架构与评测 1417 篇

2509.19955 2026-03-02 cs.IR 50%

Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach

多模态增强的联邦推荐:一种组级融合方法

Chunxu Zhang, Weipeng Zhang, Guodong Long, Zhiheng Xue, Riting Xia, Bo Yang

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出了一种多模态增强的联邦推荐方法,通过组级融合机制提升推荐系统在多模态特征整合方面的性能。

Comments Accepted at WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23132 2026-02-27 cs.IR cs.LG 50%

From Agnostic to Specific: Latent Preference Diffusion for Multi-Behavior Sequential Recommendation

从无偏到特定:基于潜在偏好的多行为序列推荐扩散模型

Ruochen Yang, Xiaodong Li, Jiawei Sheng, Jiangxia Cao, Xinkui Lin, Shen Wang, Shuang Yang, Zhaojie Liu, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, UCAS(UCAS网络安全学院)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 FatsMB通过潜在空间中的偏好生成,实现从无偏到特定的多行为序列推荐,提升推荐的多样性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22280 2026-02-27 cs.LG cs.AI 50%

Integrating Machine Learning Ensembles and Large Language Models for Heart Disease Prediction Using Voting Fusion

将机器学习集成与大型语言模型结合用于心脏病预测的投票融合

Md. Tahsin Amin, Tanim Ahmmod, Zannatul Ferdus, Talukder Naemul Hasan Naem, Ehsanul Ferdous, Arpita Bhattacharjee, Ishmam Ahmed Solaiman, Nahiyan Bin Noor

专题命中 融合架构与评测 :hybrid fusion(abstract)

AI总结 本研究通过结合机器学习集成与大型语言模型的投票融合方法,提高了心脏病预测的准确性与可靠性。

Comments 7 pages, 8 figures, (Accepted at a peer-reviewed conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20723 2026-02-27 cs.AI 50%

Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation

模态引导的图专家混合网络与熵触发路由用于多模态推荐

Ji Dai, Quan Fang, Dengsheng Cai

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tianjin University of Technology(天津理工大学)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 MAGNET通过模态引导的图专家混合网络与熵触发路由,提升多模态推荐中融合的可控性、稳定性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19569 2026-02-24 cs.CL cs.AI 50%

Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering

具有多视角融合的时序感知异构图推理用于时序问答

Wuzhenghong Wen, Bowen Zhou, Jinwen Huang, Xianjie Wu, Yuwei Sun, Su Pan, Liang Li, Jianting Liu

机构 * 1 School of Internet of Things, Nanjing University of Posts Telecommunications 2 Faculty of Medicine \& Health, University of New South Wales, Sydney, Australia 3 Dept. of Biomedical Engineering \& Biotechnology, Khalifa University, Abu Dhabi 4 Beijing Information Science Technology University 5 Shortest Path Technology Co., Ltd.

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 本文提出一种时序感知异构图推理框架,通过多视角融合提升时序问答的准确性和效率。

Comments 6pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14352 2026-02-20 cs.SI 50%

Bridging the Urban Divide: Adaptive Cross-City Learning for Disaster Sentiment Understanding

弥合城市鸿沟:面向灾害情感理解的自适应跨城学习

Zihui Ma, Yiheng Chen, Runlong Yu, Afra Izzati Kamili, Fangqi Chen, Zhaoxi Zhang, Juan Li, Yuki Miura

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出自适应跨城学习框架,通过整合行为与文本数据,提升灾害情绪理解的准确性和公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01112 2026-02-17 cs.CL 50%

EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation

EmoLoom-2B:基于词典弱监督和KV-Off评估的快速基础模型筛选用于情感分类和VAD预测

Zilin Li, Weiwei Xu, Xuanbo Lu, Zheda Liu

机构 * Zheda Liu(2 刘智达)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 EmoLoom-2B通过词典弱监督和KV-Off评估,快速筛选出适用于情感分类和VAD预测的基础模型。

Comments This paper presents an initial and self-contained study of a lightweight screening pipeline for emotion-aware language modeling, intended as a reproducible baseline and system-level design reference. This latest version corrects and updates certain personal information

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16178 2026-02-17 cs.LG stat.ML 50%

SWIFT: Mapping Sub-series with Wavelet Decomposition Improves Time Series Forecasting

SWIFT:基于小波分解的子序列映射提升时间序列预测

Wenxuan Xie, Fanpu Cao

机构 * South China University of Technology, Guangzhou, China(华南理工大学)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 SWIFT通过小波分解和轻量级结构提升长期时间序列预测的效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07969 2026-02-16 eess.AS cs.AI cs.LG cs.SD 50%

Tuberculosis Screening from Cough Audio: Baseline Models, Clinical Variables, and Uncertainty Quantification

咳嗽音频中肺结核筛查:基线模型、临床变量和不确定性量化

George P. Kafentzis, Efstratios Selisios

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出了一种标准化框架,用于通过咳嗽音频和临床数据检测肺结核,通过建立基线模型和量化不确定性,提供公平比较的参考点。

Comments Updated to published version in Sensors; DOI: 10.3390/s26041223

Journal ref Sensors 2026, 26(4), 1223

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12087 2026-02-13 cs.LG 50%

Geometry of Uncertainty: Learning Metric Spaces for Multimodal State Estimation in RL

不确定性几何:用于强化学习中多模态状态估计的度量空间学习

Alfredo Reichlin, Adriano Pacciarelli, Danica Kragic, Miguel Vasco

机构 * Division of Robotics, Perception, and Learning(机器人、感知与学习系) KTH Royal Institute of Technology(皇家理工学院)

专题命中 融合架构与评测 :sensor fusion(abstract)

AI总结 本文提出了一种学习多模态状态估计的度量空间方法,通过自适应整合多种传感器模态,提升了RL任务中的鲁棒性和状态估计性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13180 2026-02-12 cs.LG cs.AI cs.DC 50%

GC-Fed: Gradient Centralized Federated Learning with Partial Client Participation

GC-Fed: 带部分客户端参与的梯度集中联邦学习

Jungwon Seo, Ferhat Ozgur Catak, Chunming Rong, Kibeom Hong, Minhoe Kim

机构 * Department of Electrical Engineering and Computer Science, University of Stavanger(电子工程与计算机科学系,斯瓦尔内尔大学) Department of Software, Sookmyung Women's University(软件系,淑明女子大学) Department of Electrical and Information Engineering, Seoul National University of Science and Technology(电子与信息工程系,首尔科学技术大学)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 GC-Fed通过梯度集中技术缓解联邦学习中的客户端漂移问题,提升异构数据下的模型准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10023 2026-02-11 cs.CL 50%

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

MEVER:基于图的证据检索的多模态和可解释性声明验证

Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang, Dongwon Lee

机构 * University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) University of California San Diego(加州大学圣地亚哥分校) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract)

AI总结 MEVER通过多模态图检索和解释生成,实现了准确且可解释的声明验证,同时创建了AI领域的科学数据集AIChartClaim。

Comments Accepted to EACL-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14863 2026-02-04 cs.LG cs.AI 50%

Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study

探索图变换器中的全局到局部注意力机制:一项实证研究

Gang Wu, Zhengwei Wang

机构 * School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 G2LFormer通过全局到局部注意力机制提升图变换器性能,结合注意力与GNN模块,缓解信息丢失并保持线性复杂度。

Comments The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-026-51718-4}

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00561 2026-02-03 cs.AI 50%

Uncovering Latent Communication Patterns in Brain Networks via Adaptive Flow Routing

通过自适应流路由揭示脑网络中的潜在通信模式

Tianhao Huang, Guanghui Min, Zhenyu Lei, Aiying Zhang, Chen Chen

机构 * University of Virginia(弗吉尼亚大学)

专题命中 融合架构与评测 :multi-modal fusion(abstract)

AI总结 本文提出AFR-Net,通过神经通信动态视角融合SC和FC,揭示潜在神经通路,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22610 2026-02-02 cs.LG cs.AI 50%

Local-Global Multimodal Contrastive Learning for Molecular Property Prediction

局部-全局多模态对比学习用于分子性质预测

Xiayu Liu, Zhengyi Lu, Yunhong Liao, Chan Fan, Hou-biao Li

机构 * School of Mathematical Sciences(数学科学学院) University of Electronic Science and Technology of China(电子科技大学) Department of Computer Science and Engineer(计算机科学与工程系) Oakland University(奥克兰大学) Department of Electrical and Computer Engineering(电气与计算机工程系) College of Management Science(管理科学学院) Chengdu University of Technology(成都理工大学)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 LGM-CL通过局部-全局多模态对比学习框架,整合分子结构和化学语义信息,提升分子性质预测的准确性与性能。

Comments 16 pages, 9 figures. Submitted to Briefings in Bioinformatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18424 2026-01-27 cs.HC cs.LG 50%

Fusion of Spatio-Temporal and Multi-Scale Frequency Features for Dry Electrodes MI-EEG Decoding

时空与多尺度频率特征融合用于干电极MI-EEG解码

Tianyi Gong, Can Han, Junxi Wu, Dahong Qian

专题命中 融合架构与评测 :decision-level fusion(abstract)

AI总结 STGMFM通过融合时空与多尺度频率特征,提升干电极MI-EEG解码的稳定性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11885 2026-01-21 cs.AI 50%

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

MyGram: 多模态实体对齐的模态感知图变换器与全局分布

Zhifei Li, Ziyue Qin, Xiangyu Luo, Xiaoju Hou, Yue Zhao, Miao Zhang, Zhifang Huang, Kui Xiao, Bing Yang

专题命中 融合架构与评测 :multi-modal fusion(abstract)

AI总结 MyGram通过模态感知图变换器与全局分布机制,提升多模态实体对齐的性能。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10986 2026-01-19 cs.LG 50%

StellarF: A Physics-Informed LoRA Framework for Stellar Flare Forecasting with Historical & Statistical Data

StellarF:一种结合物理知识的LoRA框架,用于利用历史和统计数据进行恒星耀斑预测

Tianyu Su, Zhiqiang Zou, Qingyu Lu, Feng Zhang, Ali Luo, Xiao Kong, Min Li

机构 * School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机科学学院) Jiangsu Key Laboratory of Big Data Security and Intelligent Processing(江苏大数据安全与智能处理重点实验室) University of Chinese Academy of Sciences(中国科学院大学) CAS Key Laboratory of Optical Astronomy, National Astronomical Observatories(中国科学院国家天文台光学天文重点实验室) School of Astronomy and Space Science, University of Chinese Academy of Sciences(中国科学院大学天文与空间科学学院)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 StellarF通过结合物理知识和LoRA框架,利用历史和统计数据提升恒星耀斑预测的准确性与物理可解释性。

Comments 12 pages, 8 figures (5 main, 3 appendix), 7 tables (2 main, 5 appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08764 2026-01-14 cs.IR cs.SD eess.AS 50%

FusID: Modality-Fused Semantic IDs for Generative Music Recommendation

FusID: 多模态融合的语义ID用于生成音乐推荐

Haven Kim, Yupeng Hou, Julian McAuley

机构 * University of California San Diego(加州大学圣地亚哥分校)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 FusID通过多模态融合、表示学习和产品量化技术,解决生成音乐推荐中跨模态交互和ID冲突问题,提升推荐准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06848 2026-01-13 cs.CL 50%

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model

可解释的多模态基于方面的情感分析与依赖引导的大语言模型

Zhongzheng Wang, Yuanhe Tian, Hongzhi Wang, Yan Song

机构 * Harbin Institute of Technology(哈尔滨工程大学) Zhongguancun Academy(中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) University of Science and Technology of China(中国科学技术大学)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出了一种基于依赖引导的大语言模型的多模态情感分析方法,通过生成可解释的自然语言解释来提升情感分类的准确性。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02624 2026-01-07 cs.HC 50%

EgoLog: Ego-Centric Fine-Grained Daily Log with Ubiquitous Wearables

EgoLog:基于普及式可穿戴设备的以自我为中心的细粒度日常日志

Lixing He, Bufang Yang, Di Duan, Zhenyu Yan, Guoliang Xing

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 EgoLog通过融合音频与IMU数据及LLM推理,实现了基于普及式可穿戴设备的细粒度日常日志识别,提升了场景和活动识别的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20407 2026-01-01 cs.SD cs.AI cs.LG 50%

AUDRON: A Deep Learning Framework with Fused Acoustic Signatures for Drone Type Recognition

AUDRON:一种融合声学特征的深度学习框架用于无人机类型识别

Rajdeep Chatterjee, Sudip Chakrabarty, Trishaani Acharjee, Deepanjali Mishra

机构 * AmygdalaAI-India Lab(AmygdalaAI-India实验室)

专题命中 融合架构与评测 :feature-level fusion(abstract)

AI总结 AUDRON通过融合多种声学特征和深度学习技术,实现高准确率的无人机声学识别,适用于安全监控等场景。

Comments Presented at the 2025 IEEE 22nd India Council International Conference (INDICON). 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23824 2026-01-01 cs.LG 50%

MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling

MS-SSM:一种用于高效序列建模的多尺度状态空间模型

Mahdi Karami, Ali Behrouz, Peilin Zhong, Razvan Pascanu, Vahab Mirrokni

机构 * Google Research(谷歌研究) Google DeepMind(谷歌DeepMind)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 MS-SSM通过多尺度状态空间模型提升序列建模效率,尤其在长程和分层任务中表现优异。

Comments In Second Conference on Language Modeling (COLM) (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22102 2025-12-29 cs.LG 50%

Explainable Multimodal Regression via Information Decomposition

通过信息分解实现可解释的多模态回归

Zhaozhao Ma, Shujian Yu

机构 * Zhejiang University(浙江大学) Georgia Institute of Technology(佐治亚理工学院) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) UiT - The Arctic University of Norway(挪威北极大学)

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出基于部分信息分解的多模态回归框架,通过分解模态特定表示为唯一、冗余和协同组件,提升预测准确性和可解释性,并在多个数据集上验证其有效性。

Comments Project Page: https://github.com/zhaozhaoma/PIDReg

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21543 2025-12-29 cs.IR 50%

CEMG: Collaborative-Enhanced Multimodal Generative Recommendation

CEMG: 基于协作增强的多模态生成推荐

Yuzhen Lin, Hongyi Chen, Xuanjing Chen, Shaowen Wang, Ivonne Xu, Dongming Jiang

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 CEMG通过多模态融合层和残差量化变分自编码器,提升多模态生成推荐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21215 2025-12-25 eess.AS 50%

USE: A Unified Model for Universal Sound Separation and Extraction

USE: 一种统一的通用声音分离与提取模型

Hongyu Wang, Chenda Li, Xin Zhou, Shuai Wang, Yanmin Qian

专题命中 融合架构与评测 :multi-modal fusion(abstract)

AI总结 本文提出USE模型,通过统一框架结合声音分离与目标声音提取,提升两者性能,实验显示在SS和TSE任务中分别取得1.4 dB SDR提升和86%准确率。

Comments Accepted as an oral presentation by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21105 2025-12-25 cs.HC 50%

Volatile Organic Compounds for Stress Detection: A Scoping Review and Exploratory Feasibility Study with Low-Cost Sensors

挥发性有机化合物用于压力检测:一项综述和探索可行性研究,结合低成本传感器

Nicolai Plintz, Marcus Vetter, Dirk Ifenthaler

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文通过综述和探索性研究,证明低成本传感器可捕捉压力相关的VOC模式,揭示了VOC在情绪识别中的应用潜力及实施挑战。

Comments 13 pages, 5 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20948 2025-12-25 cs.CL cs.SD 50%

Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study

基于基础模型的神经精神障碍评估:一项涵盖全生命周期、多模态和多语言的研究

Zhongren Dong, Haotian Guo, Weixiang Xu, Huan Zhao, Zixing Zhang

机构 * College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) Shenzhen Research Institute, Hunan University(湖南大学深圳研究院) Ministry of Education Key Laboratory of Fusion Computing of Supercomputing and Artificial Intelligence, Hunan University(湖南大学教育部长春计算机与人工智能融合计算重点实验室)

专题命中 融合架构与评测 :multi-modal fusion(abstract)

AI总结 FEND提出一种多模态框架,利用多语言数据集评估神经精神障碍,揭示多模态融合在不同障碍检测中的性能差异及影响因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18661 2025-12-23 cs.AI 50%

ASTIF: Adaptive Semantic-Temporal Integration for Cryptocurrency Price Forecasting

ASTIF:自适应语义-时间整合用于加密货币价格预测

Hafiz Saif Ur Rehman, Ling Liu, Kaleem Ullah Qasim

机构 * Southwestern University of Finance and Economics(西南财经大学) Southwest Jiaotong University(西南交通大学)

专题命中 融合架构与评测 :information fusion(abstract)

AI总结 ASTIF通过自适应元学习整合语义和时间信息,提升加密货币价格预测的准确性和鲁棒性。

Comments 33 Pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11065 2025-12-17 cs.HC 50%

Immutable Explainability: Fuzzy Logic and Blockchain for Verifiable Affective AI

不可变的可解释性:模糊逻辑与区块链用于可验证的情感人工智能

Marcelo Fransoy, Alejandro Hossian, Hernán Merlino

专题命中 融合架构与评测 :multimodal fusion(abstract)

AI总结 本文提出不可变可解释性架构,结合模糊逻辑与区块链,实现情感AI的透明决策和可信审计。

详情

展开后加载摘要…

URL PDF HTML 收藏