arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-11 至 2026-02-11 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 13 篇

2512.00185 2026-02-11 cs.AI cs.LG 85%

Chunking Strategies for Multimodal AI Systems

多模态AI系统中的分块策略

Shashanka B R, Mohith Charan R, Seema Banu F

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 本文综述了多模态系统中分块策略的分类和技术分析,探讨了不同模态的数据处理方法及挑战。

Comments 50 pages, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09483 2026-02-11 cs.CV 85%

Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions

超越单个词对齐:通过令牌交互蒸馏多模态大语言模型

Lin Chen, Xiaoke Zhao, Kun Ding, Weiwei Feng, Changtao Miao, Zili Wang, Wenxuan Guo, Ying Wang, Kaiyuan Zheng, Bo Zhang, Zhe Li, Shiming Xiang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所信息与智能系统研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Align-TI通过令牌交互改进知识蒸馏,实现多模态大语言模型的高效压缩与性能提升

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15831 2026-02-11 cs.CV 85%

UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment

UniFit: 向基于多模态大语言模型引导的语义对齐的通用虚拟试衣迈进

Wei Zhang, Yeying Jin, Xin Li, Yan Zhang, Xiaofeng Cong, Cong Wang, Fengcai Qiao, zhichao Lian

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 UniFit通过多模态大语言模型引导的语义对齐模块,解决虚拟试衣中语义差距和数据稀缺问题,实现通用且高性能的试衣框架。

Comments accepted to AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09445 2026-02-11 cs.IR 82%

Personalized Parameter-Efficient Fine-Tuning of Foundation Models for Multimodal Recommendation

面向多模态推荐的个性化参数高效微调

Sunwoo Kim, Hyunjin Hwang, Kijung Shin

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(abstract)

AI总结 本文提出PerPEFT,一种面向多模态推荐的个性化参数高效微调策略,通过按兴趣分组并为每个组分配独立模块,提升推荐效果。

Comments To be published at The Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09934 2026-02-11 cs.CV 79%

VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

VersaViT: 通过任务引导优化增强MLLM视觉骨干

Yikun Liu, Yuan Liu, Shangzhe Di, Haicheng Wang, Zhongyin Zhao, Le Tian, Xiao Zhou, Jie Zhou, Jiangchao Yao, Yanfeng Wang, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University, China(上海交通大学人工智能学院) CMIC, Shanghai Jiao Tong University, China(上海交通大学计算机学院) WeChat AI, Tencent Inc., China(腾讯公司)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 VersaViT通过任务引导优化增强MLLM视觉骨干,解决其在密集预测任务中的性能问题,提升视觉任务的适应性与表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09485 2026-02-11 cs.AI 79%

Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models

连接效率与透明性:可解释的CoT压缩在多模态大推理模型中

Yizhi Wang, Linan Yue, Min-Ling Zhang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of Computer Network and Information Integration (SEU), Ministry of Education, China(计算机网络与信息集成重点实验室(SEU))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 XMCC通过强化学习优化的顺序决策过程,实现多模态大推理模型中可解释的CoT压缩,同时保持推理正确性和生成可解释的压缩解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04121 2026-02-11 cs.CV 79%

Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition

基于图的多模态和多视角对齐用于按键识别

Julia Lee Romero, Kyle Min, Subarna Tripathi, Morteza Karimzadeh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于图的多模态和多视角对齐框架,用于提高egocentric视频中按键识别的准确率,通过构建稀疏图结构并利用多模态特征提升性能。

Comments We expanded the paper and resubmitted as a separate submission to arXiv. This submission is outdated and readers can refer to arXiv:2506.01102

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09507 2026-02-11 cs.LG 78%

Towards Uniformity and Alignment for Multimodal Representation Learning

迈向多模态表示学习的统一性与对齐

Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) The Netherlands Cancer Institute(荷兰癌症研究所) The Arctic University of Norway(挪威北极大学) Singapore Management University(新加坡管理大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出了一种多模态表示学习的方法,通过分离对齐和统一性以解决冲突,提升模型在检索和生成任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09541 2026-02-11 cs.CV 74%

Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination

Scalpel: 通过混合高斯桥梁实现细粒度注意力激活流形对齐以缓解多模态幻觉

Ziqiang Shi, Rujie Liu, Shanshan Yu, Satoshi Munakata, Koichi Shirahata

机构 * Fujitsu Research & Development Center Co.,LTD.(Fujitsu 研究与开发中心有限公司) Fujitsu Limited(Fujitsu 有限公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV

AI总结 Scalpel通过高斯混合模型和熵最优传输减少多模态幻觉,实现注意力激活流形的细粒度对齐,提升视觉-语言模型的输出一致性。

Comments WACV 2026 (It was accepted in the first round, with an acceptance rate of 6%.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04397 2026-02-11 cs.CV 70%

Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration

多专家学习框架与状态空间模型用于光学和SAR图像配准

Wei Wang, Dou Quan, Ning Huyan, Chonghua Lv, Shuang Wang, Yunan Li, Licheng Jiao

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China(教育部智能感知与图像理解重点实验室) Hangzhou Institute of Technology, Xidian University(西安电子科技大学杭州学院) School of Computer Science, Xidian University(西安电子科技大学计算机学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ME-SSM框架,通过多专家学习和状态空间模型提升光学与SAR图像配准的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09066 2026-02-11 cs.LG cs.AI 57%

Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning

谱解耦与增强:一种双域对比框架用于表示学习

Jinjin Guo, Yexin Li, Zhichao Huang, Jun Fang, Zhiyuan Liu, Chao Liu, Pengzhang Liu, Qixia Jiang

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 SDE通过双域对比损失和谱增强策略,提升多模态表示学习的鲁棒性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14410 2026-02-11 eess.AS 57%

TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation

TTA:跨语言语音表示的转录、翻译与对齐

Wei Liu, Jiahong Li, Yiwen Shao, Dong Yu

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 eess.AS

AI总结 TTA通过大规模训练生成跨语言语音表示,提升语音识别与翻译任务的性能,优于Whisper模型。

Comments Accepted by ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09225 2026-02-11 cs.LG 50%

Barycentric alignment for instance-level comparison of neural representations

实例层面神经表示的重心对齐

Shreya Saha, Zoe Wanying He, Meenakshi Khosla

机构 * Department of XXX, University of YYY, Location, Country(YYY大学XXX系) School of ZZZ, Institute of WWW, Location, Country(WWW研究所ZZZ学院) Department of Electrical and Computer Engineering, UCSD(UCSD电子与计算机工程系) Cognitive Science Department, UCSD(UCSD认知科学系) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :cross-modal(abstract)

AI总结 通过实例层面的神经表示重心对齐,揭示了跨视觉和语言模型的表示收敛与发散特性,并展示了人类对齐的跨模态比较能力。

详情

展开后加载摘要…

URL PDF HTML 收藏