arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-25 至 2025-11-25 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 9 篇

2511.19257 2025-11-25 cs.CR cs.AI cs.LG 88%

Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Medusa: 跨模态可转移的对抗攻击用于多模态医疗检索增强生成

Yingjia Shang, Yi Liu, Huimin Wang, Furong Li, Wenfang Sun, Wu Chengyu, Yefeng Zheng

机构 * Westlake University(西湖大学) Heilongjiang University(黑龙江大学) City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 Medusa提出了一种针对多模态医疗检索增强生成系统的跨模态可转移对抗攻击方法,通过优化扰动和双循环策略实现高攻击成功率并抵御主流防御措施。

Comments Accepted at KDD 2026 First Cycle (full version). Authors marked with * contributed equally. Yi Liu is the lead author

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16654 2025-11-25 cs.CL 83%

Comparison of Text-Based and Image-Based Retrieval in Multimodal Retrieval Augmented Generation Large Language Model Systems

多模态检索增强生成大语言模型系统中基于文本和基于图像的检索比较

Elias Lumer, Alex Cardenas, Matt Melich, Myles Mason, Sara Dieter, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, Roberto Hernandez

机构 * PricewaterhouseCoopers U.S.(普华永道美国公司)

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

AI总结 本文比较了多模态RAG系统中基于文本和基于图像的检索方法,发现直接多模态嵌入检索在性能和准确性上优于基于LLM总结的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19380 2025-11-25 cs.CV 79%

UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval

UISearch: 基于图的多模态企业UI截图检索

Maroun Ayli, Youssef Bakouny, Tushar Sharma, Nader Jalloul, Hani Seifeddine, Rima Kilany

机构 * Center For Computer Science(计算机科学中心) Saint Joseph University of Beirut(贝鲁特圣约瑟夫大学) Faculty of Computer Science(计算机科学学院) Dalhousie University(达尔豪斯大学) Murex

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 UISearch通过基于图的结构嵌入与语义检索结合,实现多模态企业UI截图检索,提升检索准确率与效率。

Comments 12 pages, 2 figures, 3 algorithms, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18983 2025-11-25 cs.CV 79%

UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

UMCL: 单模生成多模对比学习用于跨压缩率深度伪造检测

Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UMCL通过单模生成多模对比学习,提升跨压缩率深度伪造检测的鲁棒性和准确性。

Comments 24-page manuscript accepted to IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18298 2025-11-25 cs.AI 70%

Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery

跨学科知识检索与综合:一种用于科学发现的复合AI架构

Svitlana Volkova, Peter Bautista, Avinash Hiriyanna, Gabriel Ganberg, Isabel Erickson, Zachary Klinefelter, Nick Abele, Hsien-Te Kao, Grant Engberson

机构 * Aptima, Inc.(Aptima公司)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 BioSage通过整合LLMs与RAG,利用专门代理实现跨学科知识检索与综合,提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07221 2025-11-25 cs.CV cs.AI 62%

Exploring the Use of Contrastive Language-Image Pre-Training for Human Posture Classification: Insights from Yoga Pose Analysis

探索对比语言-图像预训练在人体姿态分类中的应用:从瑜伽姿势分析获得的见解

Andrzej D. Dobrzycki, Ana M. Bernardos, Luca Bergesio, Andrzej Pomirski, Daniel Sáez-Trigueros

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究利用CLIP模型在瑜伽姿势分类中取得高准确率,展示其在人体姿态识别中的潜力,并验证其在自动化系统中的应用可行性。

Journal ref Mathematics 2024, 12(1), 76

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24466 2025-11-25 cs.CV 57%

SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking

SA-Person: 基于文本的人检索与场景感知重排序

Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen, Yang Yang, Min Cao, Mang Ye, Zhen Lei

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人中心) Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统(MAIS)) Wuhan University(武汉大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 SA-Person通过整合个体外观和全局场景上下文,提升基于文本的人检索准确性,提出ScenePerson-13W数据集和两阶段检索框架。

Comments 13 pages, 8 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18598 2025-11-25 q-bio.OT 50%

Assessing Gaze and Pointing: Human Cue Interpretation by Indian Free-Ranging Dogs in a Food Retrieval Task

评估目光与指认:印度自由放养狗在食物获取任务中的人类提示解读

Srijaya Nandi, Dipanjan Roy, Aesha Lahiri, Anamitra Roy, Anindita Bhadra

专题命中 跨模态检索 :multimodal(abstract)

AI总结 研究发现印度自由放养狗在结合指认和目光提示时能准确找到隐藏食物,但单一或冲突提示下表现无显著差异,且狗的气质影响其参与意愿和接近延迟,但不影响选择准确性。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10584 2025-11-25 cs.IR 50%

DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System

DAS: 基于双对齐语义ID的工业推荐系统

Wencai Ye, Mingjie Sun, Shaoyun Shi, Peng Wang, Wenjin Wu, Peng Jiang

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 DAS通过双对齐语义ID方法,提升推荐系统中多模态内容整合与协同信号对齐效率,有效解决信息损失与灵活性问题。

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏