AI 中文总结
研究现实世界时尚多轮图像检索问题,提出FashionAM框架直接对齐多模态对话查询与时尚图库嵌入空间,避免文本化,引入DIM-Fashion数据集,实验证明该框架优于现有方法,数据集和代码将公开。
AI 中文摘要
现实世界中的时尚搜索涉及多轮交互式检索。然而,现有的多轮检索方法基于每个交互都遵循相同属性编辑范式的限制性假设,未探索异构意图转换。此外,现有方法常依赖文本化来桥接多模态查询和视觉检索,可能丢失细粒度视觉线索。为解决这些差距,我们引入DIM-Fashion,一个由7个任务的13个时尚检索数据集构建的包含26K多轮会话的基准,具有多样化意图转换和回滚行为。我们还提出FashionAM,一个直接将多模态对话查询与时尚导向的图库嵌入空间对齐的MLLM-VLP框架,避免中间文本化。广泛实验证明了FashionAM优于现有方法。数据集和代码将在接受后公开。
英文摘要
Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches often rely on textification to bridge multimodal queries and visual retrieval, which may lose fine-grained visual cues. To address these gaps, we introduce DIM-Fashion, a benchmark of 26K multi-turn sessions constructed from 13 fashion retrieval datasets across 7 tasks, featuring diverse intent transitions and rollback behaviors. We further propose FashionAM, an MLLM-VLP framework that directly aligns multimodal conversational queries with a fashion-oriented gallery embedding space, avoiding intermediate textification. Extensive experiments demonstrate the effectiveness of FashionAM over existing approaches. The dataset and code will be made publicly available upon acceptance.