面向组合视觉数据检索的逐样本感知插值权重学习
Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval
浏览论文内容
中文总结 AI 辅助
针对组合视觉数据检索中基于MLLM方法延迟高的问题,提出SRAIN框架,通过批量排名感知权重估计和合成难负样本的内存库,实现低延迟且性能优异的检索。
中文摘要 AI 辅助
组合视觉数据检索的核心是将参考视觉输入与文本修改融合为单一查询。当前最先进的方法利用多模态大语言模型(MLLM)实现这种融合,但其复杂性会导致过高的查询时间延迟,限制了可扩展性。我们转而重新审视嵌入空间内简单线性插值的有效性,并提出SRAIN——首个动态预测查询特定插值权重的框架。关键挑战在于,插值权重的质量需通过插值后的嵌入对负样本的区分度及其与真实目标的接近度来衡量,这使得收集和预测最优权重变得困难。我们通过两项关键创新克服了这一瓶颈:训练期间的批量感知排名权重估计,以及推理期间合成难负样本的紧凑内存库。SRAIN在组合视频检索中取得最佳性能,在组合图像检索中与当前最先进方法相当,同时相比基于MLLM的替代方案大幅降低了查询时间延迟。
英文摘要
At the heart of composed visual data retrieval is the fusion of a reference visual input and a textual modification into a single query. While current state-of-the-art methods utilize multimodal large language models for this fusion, their complexity introduces prohibitive querytime latency, limiting their scalability. We instead revisit the efficacy of simple linear interpolation within an embedding space, and introduce SRAIN, the first framework that dynamically predicts query-specific interpolation weights. The key challenge lies in the fact that the quality of an interpolation weight should be measured by the interpolated embedding's discriminability from negatives as well as its proximity to true targets; this makes collecting and predicting optimal weights intractable. We overcome this bottleneck through two key innovations: batch-wise rank-aware weight estimation during training, and a compact memory bank that synthesizes hard negatives during inference. SRAIN achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval, all while substantially reducing querytime latency compared to MLLM-based alternatives.
发表机构
- AI Center, Samsung Electronics(三星电子AI中心)
- POSTECH(浦项科技大学)
- Ajou University(明知大学)
机构由 AI 辅助整理,请以论文原文为准。