MASCOT:面向复合属性文本到图像检索的模型感知子模覆盖
MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval
浏览论文内容
中文总结 AI 辅助
针对现有流形重排序方法在复合属性文本到图像检索的多样性降低任务中早期排名召回率大幅下降的问题,提出MASCOT方法,在PixelProse数据集复合约束任务中表现优于MS-DPP,尤其在排名1后召回率上优势显著。
中文摘要 AI 辅助
视觉-语言模型(VLMs)在检索语义相关图像方面效果极佳,但实际应用中仅靠相关性往往不够,系统还需实现地理、时间等复合属性层面的结果多样化(RD),而精准控制该任务仍具挑战性。当前重排序方法如多源确定性点过程(MS-DPP)通过在相似度表示上实施流形排斥来解决该问题,尽管此策略对广泛探索有效,但流形模型存在关键局限:在离散元数据的多样性降低任务中,早期排名召回率会大幅下降。为填补这一空白,本文提出MASCOT(Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval,面向复合属性文本到图像检索的模型感知子模覆盖)。MASCOT不依赖流形排斥,而是将多属性多样性表述为资源分配问题,根据查询驱动的重要性将属性投影到软分箱空间。在三项PixelProse多样性降低任务的平均值上,MASCOT保持88.58%的早期排名召回率(R@10),而MS-DPP仅为67.63%。在复合约束下差距进一步扩大:在需同时抑制时间和地理多样性的PP_geo_hour任务中,MS-DPP的召回率从0.9737降至0.4931,其排名第一的结果降至R@1=0.23,而MASCOT在多样性指标高于无约束基线的情况下,仍保持R@10=0.9410和R@1=0.7202。本文并非主张全面优势:在综合多样性-相关性得分上,本文更简单的 ablation 模型在全部三项降低任务中取得更高调和均值,MASCOT的优势仅体现在复合约束下排名1之后的召回率上。
英文摘要
Vision-Language Models (VLMs) are highly effective in retrieving semantically relevant images. However, in practice, relevance alone is often insufficient. Systems must also achieve Result Diversification (RD) across composite attributes such as geography and time, a task for which precise control remains challenging. Current re-ranking methods, such as Multi-Source Determinantal Point Processes (MS-DPP), address this using manifold-based repulsion over similarity representations. Although this strategy is effective for broad exploration, it exposes a key limitation in manifold-based models: when subjected to diversity-decrease tasks on discrete metadata, they suffer substantial degradation in early-rank recall. To bridge this gap, we introduce MASCOT (Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval). Instead of relying on manifold repulsion, MASCOT formulates multi-attribute diversity as a resource allocation problem, projecting attributes into a soft-binning space weighted by query-driven importance. Averaged across the three PixelProse diversity-decrease tasks, MASCOT preserves an early-rank recall (R@10) of 88.58%, while MS-DPP retains 67.63%. The margin widens under composite constraints: on PP_geo_hour, where temporal and geographic diversity must be suppressed simultaneously, MS-DPP's recall collapses from 0.9737 to 0.4931 and its top-ranked result degrades to R@1 = 0.23, while MASCOT holds R@10 = 0.9410 and R@1 = 0.7202 at a diversity metric above the unconstrained baseline. We do not claim uniform superiority: on aggregate diversity-relevance scores our own simpler ablations attain higher harmonic means on all three decrease tasks, and MASCOT's advantage is specific to recall beyond rank 1 under composite constraints.
发表机构
- Indian Institute of Technology Bombay(印度孟买理工学院)
机构由 AI 辅助整理,请以论文原文为准。