Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
机构 * Queen Mary University of London(伦敦女王大学) ; Samsung AI Centre(三星人工智能中心) ; Technical University of Iași(伊阿苏技术大学)
专题命中 多模态RAG :retriever(abstract)
Comments Accepted at EMNLP 2025