arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-20 至 2025-11-20 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 6 篇

2511.15435 2025-11-20 cs.CV cs.AI cs.IR 84%

HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation

HV-Attack:多模态检索增强生成的分层视觉攻击

Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan, Hanjiang Lai

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-Sen University(中山大学) Singapore Management University(新加坡管理学院)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种分层视觉攻击方法,通过在图像输入中添加不可察觉扰动,破坏多模态检索增强生成系统的检索和生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19579 2025-11-20 cs.CV cs.AI cs.LG 81%

Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration

Francisco Mena, Dino Ienco, Cassio F. Dantas, Roberto Interdonato, Andreas Dengel

机构 * Department of Computer Science, University of Kaiserslautern-Landau (RPTU)(科斯拉尔特伦大学计算机科学系) SDS, German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) INRAE, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) CIRAD, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) INRIA, EVERGREEN, University of Montpellier(蒙彼利埃大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at the Machine Learning journal, CfP: Discovery Science 2024

Journal ref Machine Learning 114, 279 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 70%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15488 2025-11-20 cs.CV cs.AI 62%

FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection

Jiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu, Xing Hu, Yan Chen, Dawei Yang

机构 * HOUMO AI

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments This paper is acceptted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08008 2025-11-20 cs.AI 57%

Combining LLM Semantic Reasoning with GNN Structural Modeling for Multi-View Multi-Label Feature Selection

Zhiqi Chen, Yuzhou Liu, Jiarui Liu, Wanfu Gao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14160 2025-11-20 cs.CL cs.LG 57%

Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * Queen Mary University of London(伦敦大学玛丽女王学院) Institut Jožef Stefan(乔泽夫·斯蒂芬研究所)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏