arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-02 至 2026-03-02 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 15 篇

2507.05394 2026-03-02 cs.CV cs.LG 83%

pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models

pFedMMA: 为视觉-语言模型设计的个性化联邦微调多模态适配器

Sajjad Ghiasvand, Mahnoosh Alizadeh, Ramtin Pedarsani

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 pFedMMA通过多模态适配器实现视觉-语言模型的个性化联邦微调,平衡个性化与泛化能力,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24027 2026-03-02 cs.CV cs.MM 81%

GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models

GuardAlign: 多模态大语言模型中的测试时安全性对齐

Xingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang, Yin Zhang, Xiang Wang, Xiangnan He

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(摩埃关键实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) Tianjin University(天津大学)

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 GuardAlign通过OT增强的安全检测和跨模态注意力校准,有效提升多模态大语言模型在测试时的安全性,减少不安全响应率并提升任务表现。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24041 2026-03-02 cs.CV 79%

Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation

仔细观察:多模态大语言模型中的自适应视觉增强以缓解幻觉

Xingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang, Zhicai Wang, Beier Zhu, Hanwang Zhang

机构 * MoE Key Lab of BIPC, University of Science and Technology of China(信息与电子技术联合实验室,中国科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出自适应视觉增强框架AIR,通过减少冗余标记和选择性整合补丁来缓解多模态大语言模型中的幻觉问题。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04576 2026-03-02 cs.CV 79%

TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification

TARDis: 用于不完整多模态肿瘤分割与分类的时间衰减表示解耦

Zishuo Wan, Qinqin Kang, Na Li, Yi Huang, Qianru Zhang, Le Lu, Yun Bian, Dawei Ding, Ke Yan

机构 * School of Automation and Electrical Engineering, University of Science and Technology Beijing(北京科技大学自动化与电气工程学院) Alibaba Group DAMO Academy(阿里巴巴集团DAMO学院) Hupan Lab(湖畔实验室) Departments of Radiology, Changhai Hospital(上海长海医院放射科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 TARDis通过时间衰减表示解耦框架,解决多模态肿瘤分割与分类中缺失模态问题,提升诊断精度并降低辐射暴露。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17966 2026-03-02 cs.IR cs.CV 79%

LLM-Enhanced Multimodal Fusion for Cross-Domain Sequential Recommendation

基于大语言模型的跨域序列推荐多模态融合

Wangyu Wu, Zhenhong Chen, Wenqiao Zhang, Xianglin Qiu, Siqi Song, Xiaowei Huang, Fei Ma, Jimin Xiao

机构 * Xi'an Jiaotong-Liverpool University(西安交通大学利物浦大学) University of Liverpool(利物浦大学) Microsoft(微软公司) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出LLM-EMF方法,通过融合视觉和文本数据提升跨域序列推荐性能,利用CLIP模型生成多模态嵌入并引入多重注意力机制,实验证明其在多领域推荐中的有效性。

Comments arXiv admin note: substantial text overlap with arXiv:2504.15085

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01728 2026-03-02 cs.CV 79%

Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion

Shuffle Mamba:基于随机洗牌的态空间模型用于多模态图像融合

Ke Cao, Xuanhua He, Tao Hu, Chengjun Xie, Man Zhou, Jie Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Intelligent Machines(智能机器研究所) Hefei Institutes of Physical Science, Chinese Academy of Sciences(中国科学院合肥物质科学研究院) Intelligent Agriculture Engineering Laboratory of Anhui Province, Institute of Intelligent Machines(安徽省智能农业工程实验室,智能机器研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 Shuffle Mamba通过引入随机洗牌策略和逆洗牌,解决多模态图像融合中固定扫描策略带来的偏见问题,提升融合质量。

Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23862 2026-03-02 cs.HC 78%

Human-Centered Multimodal Fusion for Sexism Detection in Memes with Eye-Tracking, Heart Rate, and EEG Signals

以人类为中心的多模态融合用于表情包中性别歧视检测:结合眼动追踪、心率和EEG信号

Iván Arcos, Paolo Rosso, Elena Gomis-Vicent

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种结合眼动追踪、心率和EEG信号的多模态融合模型,用于更准确地检测表情包中的性别歧视。

Comments Accepted to appear in the Proceedings of LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19955 2026-03-02 cs.IR 78%

Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach

多模态增强的联邦推荐:一种组级融合方法

Chunxu Zhang, Weipeng Zhang, Guodong Long, Zhiheng Xue, Riting Xia, Bo Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种多模态增强的联邦推荐方法,通过组级融合机制提升推荐系统在多模态特征整合方面的性能。

Comments Accepted at WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23699 2026-03-02 cs.CV cs.CL 73%

HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit

HiDrop:通过晚期注入、凹形金字塔剪枝和早期退出实现MLLM中的层次视觉令牌减少

Hao Wu, Yingqi Fan, Jinyang Dai, Junlong Tong, Yunpu Ma, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Munich Center for Machine Learning, LMU Munich(慕尼黑大学机器学习中心,慕尼黑大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 HiDrop通过晚期注入、凹形金字塔剪枝和早期退出机制,实现多模态大语言模型中视觉令牌的高效减少,提升训练效率并保持性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23959 2026-03-02 cs.CV 70%

Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought

通过图像作为连续动作进行思考:数值视觉链式推理

Kesen Zhao, Beier Zhu, Junbao Zhou, Xingyu Zhu, Zhongqi Yue, Hanwang Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Science and Technology of China(中国科学技术大学) Chalmers University of Technology(楚克理工大学) University of Gothenburg(哥德堡大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 NV-CoT通过将图像推理动作空间扩展为连续欧几里得空间,提升MLLMs的定位精度和回答准确性,同时加速训练收敛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23734 2026-03-02 cs.CV cs.CL 62%

UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking

UTPTrack: 向简单和统一的令牌剪枝迈进:视觉跟踪中的统一令牌剪枝

Hao Wu, Xudong Wang, Jialiang Zhang, Junlong Tong, Xinghao Chen, Junyan Lin, Yunpu Ma, Xiaoyu Shen

机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学) Munich Center for Machine Learning, LMU Munich(慕尼黑机器学习中心,慕尼黑大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 UTPTrack通过统一令牌剪枝框架,在视觉跟踪中实现高精度与高效能的平衡,同时支持多模态和语言引导任务。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03059 2026-03-02 cs.CV 57%

CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning

CLAP:通过曲率采样和原型学习实现融合3D感知的无监督3D表示学习

Runjian Chen, Hang Zhang, Avinash Ravichandran, Hyoungseob Park, Wenqi Shao, Alex Wong, Ping Luo

机构 * The University of Hong Kong(香港大学) Cruise Yale University(耶鲁大学) Shanghai AI Laboratory(上海人工智能实验室) HKU Shanghai Intelligent Computing Research Center(香港大学上海智能计算研究中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 CLAP通过曲率采样和原型学习,在无监督3D表示学习中实现图像与点云的联合预训练,提升融合3D感知的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22294 2026-03-02 cs.LG 50%

When Should a Model Change Its Mind? An Energy-Based Theory and Regularizer for Concept Drift in Electrocardiogram (ECG) Signals

模型何时应改变观点?一种基于能量的理论和正则化器用于心电图(ECG)信号中的概念漂移

Timothy Oladunni, Blessing Ojeme, Kyndal Maclin, Clyde Baidoo

机构 * Department of Computer Science, Morgan State University(计算机科学系,莫根州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本研究提出Physiologic Energy Conservation Theory (PECT)和Energy-Constrained Representation Learning (ECRL),用于在ECG信号中稳定概念,通过能量守恒原理减少误判和漂移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18579 2026-03-02 cs.LG 50%

Sparsity Forcing: Reinforcing Token Sparsity of MLLMs

稀疏性强制:强化多模态大语言模型的token稀疏性

Feng Chen, Yefei He, Lequan Lin, Chenhui Gou, Jing Liu, Bohan Zhuang, Qi Wu

机构 * AIML, University of Adelaide, Australia(AIML,澳大利亚阿德莱德大学) ZIP Lab, Zhejiang University, China(浙江工业大学ZIP实验室) University of Sydney, Australia(悉尼大学) Monash University, Australia(墨尔本大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出Sparsity Forcing方法,通过强化学习框架在多模态大语言模型中提升token稀疏性,实现75%的token减少与极小的精度损失。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05629 2026-03-02 cs.LG 50%

On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

在SFT的泛化上:从强化学习视角的奖励校正

Yongliang Wu, Yizhou Zhou, Zhou Ziheng, Yingzhe Peng, Xinyu Ye, Xinting Hu, Wenbo Zhu, Lu Qi, Ming-Hsuan Yang, Xu Yang

机构 * Southeast University(东南大学) University of California, Los Angeles(加州大学洛杉矶分校) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) University of California, Berkeley(加州大学伯克利分校) Wuhan University(武汉大学) University of California, Merced(加州大学默塞德分校)

专题命中 多模态训练与对齐 :multi-modal(abstract)

AI总结 本文提出动态微调方法,通过奖励校正提升SFT的泛化能力,在多个任务中表现优于传统SFT。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏