arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-14 至 2026-01-14 共收录 14 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 14 篇

2508.05097 2026-01-14 cs.CL 83%

MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification

MultiCheck: 通过统一多模态事实验证加强网络信任

Aditya Kishore, Gaurav Kumar, Jasabanta Patro

机构 * IISER Bhopal(比哈尔IISER)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 MultiCheck通过统一多模态事实验证框架,提升网络信任,具备高效、透明和抗噪能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08399 2026-01-14 cs.LG cs.AI cs.CV 81%

Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment

通过不一致来对齐:面向多模态对齐的边界感知课程学习

Hua Ye, Hang Ding, Siyuan Chen, Yiyang Jiang, Changyuan Zhang, Xuan Zhang

机构 * Nanjing University(南京大学) Airon Technology CO., LTD(艾润科技有限公司) Shanghai Jiao Tong University(上海交通大学) University of Bristol(布里斯托大学) The Hong Kong Polytechnic University(香港理工大学) The University of Hong Kong(香港大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出BACL方法,通过边界感知负样本采样和局部注意力损失,提升多模态对齐性能,在多个基准上取得优于CLIP的成果。

Comments 24 pages, 6 figures, 5 tables. Submitted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08503 2026-01-14 cs.LG cs.AI 79%

Temporal Fusion Nexus: A task-agnostic multi-modal embedding model for clinical narratives and irregular time series in post-kidney transplant care

时间融合 nexus:一种任务无关的多模态嵌入模型,用于术后肾移植护理中的临床叙述和不规则时间序列

Aditya Kumar, Simon Rauch, Mario Cypko, Marcel Naik, Matthieu-P Schapranow, Aadil Rashid, Fabian Halleck, Bilgin Osmanodja, Roland Roller, Lars Pape, Klemens Budde, Mario Schiffer, Oliver Amft

机构 * Hahn-Schickard(哈恩-施克特德研究所) University of Freiburg(弗赖堡大学) Charité University Medical Center(查理医院大学医学中心) Hasso Plattner Institute for Digital Engineering, University of Potsdam(哈索·platner数字工程研究所,波茨坦大学) DFKI(德意志联邦防务研究院) University Hospital Essen(埃森大学医院) University Hospital Erlangen(埃尔兰根大学医院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 TFN是一种多模态嵌入模型,通过整合临床文本和不规则时间序列数据,在术后肾移植护理中提高了移植物丢失、排斥和死亡预测的性能。

Comments 31 pages, 9 figures, 3 tables. A supplementary file is also available

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07856 2026-01-14 quant-ph cs.AI cs.LG 79%

Feature Entanglement-based Quantum Multimodal Fusion Neural Network

基于特征纠缠的量子多模态融合神经网络

Yu Wu, Qianli Zhou, Jie Geng, Xinyang Deng, Wen Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于特征纠缠的量子多模态融合神经网络,通过量子计算框架解决多模态学习中的精度、可解释性和复杂性矛盾,实现高效且可解释的多模态融合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07877 2026-01-14 cs.LG cs.AI 77%

E^2-LLM: Bridging Neural Signals and Interpretable Affective Analysis

E²-LLM:连接神经信号与可解释的情感分析

Fei Ma, Han Lin, Yifan Xie, Hongwei Ren, Xiaoyu Shen, Wenbo Ding, Qi Tian

机构 * Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室) Zhejiang University(浙江大学) Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学) Eastern Institute of Technology(东方技术研究所) Huawei(华为)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

AI总结 E²-LLM通过整合预训练EEG编码器与Qwen-based LLM,实现了可解释的情绪分析,展示了模型扩展在情感识别和可解释性上的优势。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08017 2026-01-14 cs.CV cs.AI 73%

Representations of Text and Images Align From Layer One

文本和图像的表示从第一层对齐

Evžen Wybitul, Javier Rando, Florian Tramèr, Stanislav Fort

机构 * D-INFK, ETH Zurich, Switzerland(苏黎世联邦理工学院信息与知识系统研究所) Aisle Research(Aisle研究)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 该研究通过合成方法证明,视觉-语言模型中图像和文本表示在第一层即可实现对齐,为模型可解释性提供了新路径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16665 2026-01-14 cs.CV 70%

A Diff-Attention Aware State Space Fusion Model for Remote Sensing Classification

一种面向遥感分类的差分注意力感知状态空间融合模型

Wenping Ma, Boyou Xue, Mengru Ma, Chuang Chen, Hekai Zhang, Hao Zhu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出DAS2F-Model,通过差分注意力模块和线性融合模块提升多模态遥感图像分类性能。

Comments After a careful review, we discovered that there were data errors in the paper, which led to the invalidity of the conclusion. To avoid misleading the readers, we have decided to withdraw this article. We appreciate your understanding and support for our work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08811 2026-01-14 cs.CV cs.AI 62%

Reasoning Matters for 3D Visual Grounding

推理在3D视觉定位中至关重要

Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai, Cheng-Yen Yang, Jen-Hao Cheng, Jenq-Neng Hwang

机构 * University of Washington(华盛顿大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种自动合成3D视觉定位数据的管道,并引入了在仅使用1.6%训练数据下表现优于现有方法的Reason3DVG-8B模型,证明了推理在3D视觉定位中的重要性。

Comments 2025 CVPR Workshop on 3D-LLM/VLA: Bridging Language, Vision and Action in 3D Environments

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08684 2026-01-14 cs.AI cs.CV 62%

MEMEWEAVER: Inter-Meme Graph Reasoning for Sexism and Misogyny Detection

MEMEWEAVER:跨迷因图推理用于性别歧视和性别歧视检测

Paolo Italiani, David Gimeno-Gomez, Luca Ragazzi, Gianluca Moro, Paolo Rosso

机构 * Department of Computer Science and Engineering, University of Bologna(博洛尼亚大学计算机科学与工程系) PRHLT Research Center, Universtitat Politècnica de València(巴塞罗那理工大学PRHLT研究中心) ValgrAI - Valencian Graduate School and Research Network of Artificial Intelligence, Spain(西班牙瓦伦西亚人工智能研究生院与研究网络)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MemeWeaver通过跨迷因图推理机制,有效检测性别歧视和性别歧视,优于现有基线方法。

Comments Accepted at EACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07871 2026-01-14 q-bio.QM cs.AI cs.CV cs.LG 62%

Imaging-anchored Multiomics in Cardiovascular Disease: Integrating Cardiac Imaging, Bulk, Single-cell, and Spatial Transcriptomics

心血管疾病中的成像锚定多组学:整合心脏成像、批量、单细胞和空间转录组学

Minh H. N. Le, Tuan Vinh, Thanh-Huy Nguyen, Tao Li, Bao Quang Gia Le, Han H. Huynh, Monika Raj, Carl Yang, Min Xu, Nguyen Quoc Khanh Le

机构 * International Ph.D. Program in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan AIBioMed Research Group, Taipei Medical University, Taipei, Taiwan Medical Sciences Division, University of Oxford, Oxford, United Kingdom Computational Biology Department, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA Department of Computer Science, Emory University, Atlanta, GA, USA Department of Chemistry, Emory University, Atlanta, GA, USA International Master Program for Translational Science, College of Medical Science Technology, Taipei Medical University, Taipei 110, Taiwan In-Service Master Program in Artificial Intelligence in Medicine, College of Medicine, Taipei Medical University, Taipei, Taiwan Translational Imaging Research Center, Taipei Medical University Hospital, Taipei, Taiwan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过整合心脏成像与多组学数据,推动心血管疾病研究的多模态融合方法,提升疾病诊断和治疗的精准性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08764 2026-01-14 cs.IR cs.SD eess.AS 57%

FusID: Modality-Fused Semantic IDs for Generative Music Recommendation

FusID: 多模态融合的语义ID用于生成音乐推荐

Haven Kim, Yupeng Hou, Julian McAuley

机构 * University of California San Diego(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 eess.AS

AI总结 FusID通过多模态融合、表示学习和产品量化技术,解决生成音乐推荐中跨模态交互和ID冲突问题,提升推荐准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08619 2026-01-14 cs.CV 57%

CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion

CtrlFuse: 基于掩码提示的可控红外与可见图像融合

Yiming Sun, Yuan Ruan, Qinghua Hu, Pengfei Zhu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 CtrlFuse通过基于掩码提示的可控融合框架,实现红外与可见图像的交互式动态融合,提升任务性能与融合质量。

Comments 18 pages,22 figures,published to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08336 2026-01-14 cs.CV 57%

Tissue Classification and Whole-Slide Images Analysis via Modeling of the Tumor Microenvironment and Biological Pathways

通过建模肿瘤微环境和生物通路进行组织分类和全片图像分析

Junzhuo Liu, Xuemei Du, Daniel Reisenbuchler, Ye Chen, Markus Eckstein, Christian Matek, Friedrich Feuerhake, Dorit Merhof

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 BioMorphNet通过整合肿瘤微环境和生物通路信息,提升全片图像中组织分类和差异基因分析的准确性。

Comments 19 pages, 8 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06400 2026-01-14 cs.CV 57%

Perceptual Region-Driven Infrared-Visible Co-Fusion for Extreme Scene Enhancement

感知驱动的红外可见联合融合用于极端场景增强

Jing Tao, Yonghong Zong, Banglei Guan, Pengju Sun, Taihang Lei, Yang Shanga, Qifeng Yu

机构 * College of Aerospace Science and Engineering, National University of Defense Technology(航天科学与工程学院,国防科技大学) Hunan Provincial Key Laboratory of Image Measurement and Vision Navigation(湖南省图像测量与视觉导航重点实验室) Beijing Institute of Tracking and Telecommunication Technology(北京跟踪与电信技术研究所) National Key Laboratory of Space Integrated Information System(空间一体化信息系统国家重点实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种基于区域感知的红外可见联合融合框架,通过多曝光和多模态成像技术,在极端条件下提升图像清晰度和融合性能。

Comments The paper has been accepted and officially published by OPTICS AND LASER TECHNOLOGY

详情

展开后加载摘要…

URL PDF HTML 收藏