arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-27 至 2026-02-27 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2602.22678 2026-02-27 cs.CV cs.AI 84%

ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport

ViCLIP-OT:首个面向越南语图像-文本检索的视觉-语言基础模型,结合最优传输

Quoc-Khang Tran, Minh-Thien Nguyen, Nguyen-Khang Pham

机构 * Can Tho University(金兰大学)

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 ViCLIP-OT通过整合CLIP对比学习与SIGROT损失,提升越南语图像-文本检索性能,实现域内和零样本设置下的显著改进。

Comments Preprint submitted to Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23229 2026-02-27 cs.CV 79%

Large Multimodal Models as General In-Context Classifiers

大多模态模型作为通用上下文分类器

Marco Garosi, Matteo Farina, Alessandro Conti, Massimiliano Mancini, Elisa Ricci

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CIRCLE方法,通过伪标签迭代优化,使LMM在开放世界分类中超越VLM,展示LMM作为统一分类器的潜力。

Comments CVPR Findings 2026. Project website at https://circle-lmm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22734 2026-02-27 cs.CV 79%

Asymmetric Idiosyncrasies in Multimodal Models

多模态模型中的非对称个性化特征

Muzi Tao, Chufan Shi, Huijuan Wang, Shengbang Tong, Xuezhe Ma

机构 * University of Southern California(南加州大学) New York University(纽约大学)

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV

AI总结 本文研究了多模态模型中描述模型的个性化特征及其对文本到图像模型的影响,发现生成图像丢失了描述中的关键变化,提出了一种新的量化方法。

Comments Project page: https://muzi-tao.github.io/asymmetric-idiosyncrasies/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22644 2026-02-27 cs.CV 79%

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

插件、即插即用并加固:一种低成本模块用于鲁棒多模态图像理解模型

Siqi Lu, Wanying Xu, Yongbin Zheng, Wenting Luan, Peng Sun, Jianhang Yao

机构 * National University of Defense Technology(国防科技大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种低成本模块,通过频域分析解决多模态模型中缺失模态导致的性能问题,提升模型鲁棒性和整体学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11221 2026-02-27 cs.CL 79%

The Automatic Verification of Image-Text Claims (AVerImaTeC) Shared Task

图像-文本主张的自动验证(AVerImaTeC)共享任务

Rui Cao, Zhenyun Deng, Yulong Chen, Michael Schlichtkrull, Andreas Vlachos

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CL

AI总结 本文提出图像-文本主张自动验证共享任务,通过检索证据和验证真实主张,评估系统性能并展示最佳结果。

Comments Shared Task Overview and Summary for the Ninth FEVER Workshop, Co-located at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20570 2026-02-27 cs.CV cs.AI 62%

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

Dyslexify: 一种针对CLIP中印刷攻击的机制性防御

Lorenz Hufe, Constantin Venhoff, Erblina Purelku, Maximilian Dreyer, Sebastian Lapuschkin, Wojciech Samek

机构 * Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所) University of Oxford(牛津大学) Technological University Dublin(都柏林技术大学) Technische Universität Berlin(柏林技术大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 Dyslexify通过消融CLIP中的印刷电路,有效防御印刷攻击,提升性能并保持应用安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05435 2026-02-27 eess.AS cs.AI cs.LG 62%

Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning

无偏切片Wasserstein核用于高质量音频描述生成

Manh Luong, Khai Nguyen, Dinh Phung, Gholamreza Haffari, Lizhen Qu

机构 * Monash University(墨尔本大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出无偏切片Wasserstein核,通过保留模态间时间信息,提升音频描述质量及推理能力。

Journal ref Manh Luong. (2025). Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning. In Advances in Neural Information Processing Systems 38 (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18897 2026-02-27 cs.CV 57%

Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs

超越标签:基于推理增强的LMMs的无词汇细粒度识别

Dmitry Demidov, Zaigham Zaheer, Zongyan Han, Omkar Thawakar, Rao Anwer

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 FiNDR通过推理增强的LMMs实现无词汇细粒度识别,提出三步流程生成候选标签、过滤排名并构建轻量级分类器,取得显著性能提升。

Journal ref CVPR 2026 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05535 2026-02-27 cs.LG 50%

Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification

通过证据不确定性量化检测大视觉-语言模型的误行

Tao Huang, Rui Wang, Xiaofei Liu, Yi Qin, Li Duan, Liping Jing

机构 * State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院) School of Automation and Intelligence, Beijing Jiaotong University(北京交通大学自动化与智能学院) Beijing Key Laboratory of Security and Privacy in Intelligent Transportation(北京智能交通安全与隐私重点实验室)

专题命中 图文多模态 :multimodal(abstract)

AI总结 通过证据不确定性量化检测大视觉-语言模型的误行,识别内部冲突和无知以提高模型可靠性。

Comments Accepted to ICLR 2026. Code is available at https://github.com/HT86159/EUQ

详情

展开后加载摘要…

URL PDF HTML 收藏