arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-04 至 2026-02-04 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2406.15098 2026-02-04 cs.LG cs.AI 83%

How Intermodal Interaction Affects the Performance of Deep Multimodal Fusion for Mixed-Type Time Series

多模态交互如何影响深度多模态融合在混合类型时间序列中的性能

Simon Dietz, Thomas Altstidl, Dario Zanca, Björn Eskofier, An Nguyen

机构 * FAU Erlangen-Nürnberg(弗赖堡大学埃尔朗根-纽伦堡分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文研究了多模态交互对深度多模态融合在混合类型时间序列预测中的影响,通过三种融合类型和五种融合方法的比较,揭示了交互强度和方向对融合策略选择的关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02784 2026-02-04 cs.LG cs.AI 83%

Cross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning

跨时间注意力融合(CTAF)用于自监督学习中的多模态生理信号

Arian Khorasani, Théophile Demazure

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 CTAF通过时间感知的交叉注意力机制,实现多模态生理信号在自监督学习中的高效融合,提升匹配对的余弦边距和跨模态检索性能,同时保持高准确率并减少标签依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03622 2026-02-04 cs.CV physics.med-ph 80%

Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis

基于准多模态的病理特征学习用于视网膜疾病诊断

Lu Zhang, Huizhen Yu, Zuowei Wang, Fu Gui, Yatu Guo, Wei Zhang, Mengyu Jia

机构 * Tianjin University(天津大学) Tianjin Key Laboratory of Ophthalmology and Visual Science(天津眼科学与视觉科学重点实验室) Tianjin Eye Institute(天津眼科研究院) Tianjin Eye Hospital(天津眼科医院) Clinical College of Ophthalmology, Tianjin Medical University(天津医科大学临床医学院) Department of Ophthalmology, The Second Affiliated Hospital of Nanchang University(南昌大学第二附属医院眼科部) Nankai University Affiliated Eye Hospital(南开大学附属眼科医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于准多模态的视网膜疾病诊断方法,通过多模态数据合成与融合提升分类和分级的准确性。

Journal ref Zhang, L., Yu, H., Wang, Z., Gui, F., Guo, Y., Zhang, W., Jia, M., 2026. Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis. Medical Image Analysis 109, 103886

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03815 2026-02-04 cs.CV cs.LG 79%

Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning

通过视觉标记修剪实现多模态大语言模型的快慢高效训练

Dingkun Zhang, Shuhan Qi, Yulin Wu, Xinyu Xiao, Xuan Wang, Long Chen

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 DualSpeed通过快慢双模式实现多模态大语言模型的高效训练,提升训练速度并保持性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03665 2026-02-04 cs.CV cs.HC 79%

MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment

MM-SCALE: 基于标量判断和列表对齐的 grounded 多模态道德推理

Eunkyu Park, Wesley Hanwen Deng, Cheyon Jin, Matheus Kunzler Maldaner, Jordan Wheeler, Jason I. Hong, Hong Shen, Adam Perer, Ken Holstein, Motahhare Eslami, Gunhee Kim

机构 * Seoul National University(首尔国立大学) Carnegie Mellon University(卡内基梅隆大学) University of Florida(佛罗里达大学) Epic Games

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MM-SCALE通过5点标量评分和显式模态接地,提升多模态模型在道德推理任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11442 2026-02-04 cs.CV 79%

MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder

MultiMAE用于脑部MRI:通过多模态掩码自编码器提高对缺失输入的鲁棒性

Ayhan Can Erdur, Christian Beischl, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C Peeken

机构 * Department of Radiation Oncology, TUM University Hospital, Munich, Germany(辐射肿瘤科,技术大学慕尼黑医院,慕尼黑,德国) Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM)(医疗与医学人工智能教研室,技术大学慕尼黑(TUM)) TUM University Hospital, Munich, Germany(技术大学慕尼黑医院,慕尼黑,德国) Chair for AI for Image-Guided Diagnosis and Therapy, Technical University of Munich (TUM)(图像引导诊断与治疗人工智能教研室,技术大学慕尼黑(TUM)) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心(MCML),慕尼黑,德国) Department of Computing, Imperial College London, London, UK(计算系,伦敦帝国学院,伦敦,英国) Deutsches Konsortium für Translationale Krebsforschung (DKTK), Partner Site Munich, Munich, Germany(德国转化癌症研究联盟(DKTK),慕尼黑合作伙伴站点,慕尼黑,德国) Institute of Radiation Medicine (IRM), Department of Radiation Sciences (DRS), Helmholtz Center Munich, Munich, Germany(辐射医学研究所(IRM),辐射科学部门(DRS),慕尼黑海德堡中心,慕尼黑,德国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 MultiMAE通过多模态掩码自编码器提升脑部MRI在缺失输入下的鲁棒性,实现分割和分类任务的性能提升。

Comments Official implementation: https://github.com/chris-beischl/multimae-for-brain-mri

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18742 2026-02-04 cs.LG 78%

Constraint Matters: Multi-Modal Representation for Reducing Mixed-Integer Linear programming

约束至关重要:用于减少混合整数线性规划的多模态表示

Jiajun Li, Yixuan Li, Ran Hou, Yu Ding, Shisi Guan, Jiahui Duan, Xiongwei Han, Tao Zhong, Vincent Chau, Weiwei Wu, Wanyuan Wang

机构 * Southeast University(东南大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

AI总结 本文提出了一种基于约束的多模态表示方法,用于提高混合整数线性规划问题的求解效率和解的质量。

Comments Accecpted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18533 2026-02-04 cs.CV cs.CL cs.CR 62%

Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models

重新审视视觉语言模型安全微调中的瓶颈

Yi Ding, Lijun Li, Bing Cao, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Purdue University(普渡大学) Tianjin University(天津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出多图像安全数据集,通过增强视觉推理能力,提升模型在安全关键任务中的性能与通用能力。

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18179 2026-02-04 cs.CL cs.AI 62%

Problem Solved? Information Extraction Design Space for Layout-Rich Documents using LLMs

问题已解决?利用LLMs处理布局丰富文档的信息提取设计空间

Gaye Colakoglu, Gürkan Solmaz, Jonathan Fürst

机构 * Zurich University of Applied Sciences(苏黎世应用科学大学) NEC Laboratories Europe(NEC欧洲实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文通过LayIE-LLM测试套件研究了利用LLMs处理布局丰富文档的信息提取设计空间,证明通用LLMs在优化配置下可媲美专用模型,提供低成本无微调方案。

Comments accepted at EMNLP'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09125 2026-02-04 cs.CV cs.AI 62%

HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition

HAAP: 基于自适应排列的视觉-上下文分层注意力自回归模型

Honghui Chen, Yuhang Qiu, Jiabao Wang, Pingping Chen, Nam Ling

机构 * College of Physics and Information Engineering, Fuzhou University(福州大学物理与信息工程学院) Faculty of Engineering, Monash University(莫纳什大学工程学院) Department of Computer Science and Engineering, Santa Clara University(圣克拉拉大学计算机科学与工程系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 HAAP通过自适应排列和跨模态分层注意力机制,提升场景文本识别的准确性和效率。

Comments 12 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03530 2026-02-04 cs.CV 57%

Interpretable Logical Anomaly Classification via Constraint Decomposition and Instruction Fine-Tuning

通过约束分解和指令微调实现可解释的逻辑异常分类

Xufei Zhang, Xinjiao Zhou, Ziling Deng, Dongdong Geng, Jianxiong Wang

机构 * Beijing XingYun Digital Technology Co., Ltd.(北京星云数字技术有限公司)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

AI总结 LogiCls通过约束分解和指令微调实现工业逻辑异常的可解释分类,结合视觉-语言框架和数据驱动的监督方法,提升异常检测的准确性和可解释性。

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19154 2026-02-04 cs.CV 57%

Lightweight RGB-T Tracking with Mobile Vision Transformers

轻量级RGB-T跟踪与Mobile Vision Transformers

Mahdi Falaki, Maria A. Amer

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Concordia University(康科迪亚大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于MobileViT的轻量级RGB-T跟踪器,通过渐进融合框架和可分离混合注意力模型,实现低参数量和高实时性能的多模态跟踪。

Comments Accepted for publication in ICASSP 2026. Implementation Code Available

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02546 2026-02-04 cs.LG cs.AI 57%

GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning

GPG:一种简单而强大的用于模型推理的强化学习基线

Xiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei, Yong Wang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 GPG提出了一种简化且高效的强化学习方法,通过直接优化原始目标并消除冗余组件,实现了在多种任务中优于传统方法的性能。

Comments Accepted to ICLR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03319 2026-02-04 cs.LG cond-mat.mtrl-sci cs.IT math.IT 50%

Information-Theoretic Multi-Model Fusion for Target-Oriented Adaptive Sampling in Materials Design

信息论多模型融合用于材料设计中的目标导向自适应采样

Yixuan Zhang, Zhiyuan Li, Weijia He, Mian Dai, Chen Shen, Teng Long, Hongbin Zhang

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出了一种基于信息论的多模型融合框架,用于材料设计中的目标导向自适应采样,通过低熵信息状态和多模型融合提高样本效率和可靠性。

Comments 37 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00514 2026-02-04 cs.RO 50%

A Low-Cost Vision-Based Tactile Gripper with Pretraining Learning for Contact-Rich Manipulation

一种低成本的基于视觉的触觉夹具,用于接触丰富的操作

Yaohua Liu, Binkai Ou, Zicheng Qiu, Ce Hao, Hengjun Zhang

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Innovation and Research and Development Department, BoardWare Information System Company Ltd.(创新与研发部,BoardWare信息系统有限公司) School of Artificial Intelligence, Nanjing Agricultural University(人工智能学院,南京农业大学) School of Computing, National University of Singapore(计算机学院,新加坡国立大学) School of Electronic Engineering and Automation, Guilin University of Electronic Technology(电子工程与自动化学院,桂林电子科技大学)

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 本研究提出了一种低成本的视觉-触觉夹具,通过融合视觉和触觉反馈,提升接触丰富环境中的操作性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16936 2026-02-04 cs.LG 50%

SPAR: Self-supervised Placement-Aware Representation Learning for Distributed Sensing

SPAR:用于分布式传感的自监督位置感知表示学习

Yizhuo Chen, Tianchen Wang, You Lyu, Yanlan Hu, Jinyang Li, Tomoyoshi Kimura, Hongjue Zhao, Yigong Hu, Denizhan Kara, Tarek Abdelzaher

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 SPAR通过信号与位置的二元性原则,提出了一种自监督位置感知表示学习框架,提升分布式传感在多种模态和任务中的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏