arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-08-31 至 2026-08-31 共收录 2 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 2 篇

2608.20756 2026-08-31 cs.CV cs.AI 版本更新 89%

Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

Vis-Poison:多模态检索增强生成中的视觉知识投毒攻击

Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao

机构 * Southwestern University of Finance and Economics(西南财经大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 跨模态检索 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出Vis-Poison攻击,通过自动化多智能体方法构建视觉合理的被投毒图像,在黑盒设置下对多模态RAG系统实现40.16%-65.40%的攻击成功率,且对仅依赖参数知识的MLLM平均成功率超60%。

Comments Findings of EMNLP, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27865 2026-08-31 cs.PF 新提交 78%

FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval

FFSlim:一种用于多模态数据存储与检索的高效轻量格式

Long Yang, Yu Mao, Yuchen Shao, Yumiao Zhao, Yaqi Li, Xuan Liu, Xiaolong Shen, Tao Yu, Gezi Li, Jing Wang, Chengcheng Wan, Liang Shi

专题命中 跨模态检索 :multi-modal(title,abstract)

AI总结 针对多模态数据存储检索的I/O瓶颈问题,提出轻量格式FFSlim,通过三个组件提升效率,实现高吞吐量并降低端到端训练时间

详情

展开后加载摘要…

URL PDF HTML 收藏