用于可控时尚检索的属性条件多模态槽分解
Attribute-Conditioned Multimodal Slot Factorization for Controllable Fashion Retrieval
- Walmart Global Tech(沃尔玛全球科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出MM-slotgate多模态槽编码器,将Fashion-CLIP嵌入分解为属性槽,在H&M数据集上实现可控时尚检索,颜色等属性性能提升显著,优于多模态融合与仅文本检索方法。
AI中文摘要:
时尚检索通常需要同时满足多个属性,如类别、颜色、图案和人口统计特征。整体嵌入将这些信号混合到单个向量中,使得检索时难以实现属性特定的控制。许多现有的语义ID方法提供离散的物品编码,但这些编码通常被优化为物品级或残差地址,且不暴露可独立控制的命名属性槽。我们引入MM-slotgate,一种多模态槽编码器,它将Fashion-CLIP的文本和图像嵌入分解为四个命名属性槽。每个槽学习自身的文本-图像门,因此颜色和图案等视觉基础属性可更多依赖图像证据,而类别和人口统计等分类导向属性可更多保持文本驱动。在H&M数据集上,使用结合槽相似度和槽逻辑的检索分数,MM-slotgate实现了0.7566的macro ConstraintSatisfied@10,优于等权重多模态融合(0.7142)和仅文本的fCLIP检索(0.4755)。最大的提升出现在颜色上,从0.321提升至0.889(绝对提升0.568),因为学习到的颜色门为图像证据分配了57.4%的权重。这些学习到的门无需模态监督即可解释:颜色偏向图像,类别偏向文本,图案和人口统计则接近中间。生成的槽仍保持可控性:线性探测显示无超过标签相关基线的额外泄漏,且量化槽编码支持目标干预,包括颜色提升15.3倍。这些结果表明,可控时尚检索受益于类型化、属性条件的多模态槽,而非单个全局嵌入或不透明的物品级语义ID。
英文摘要:
Fashion retrieval often requires satisfying multiple attributes at once, such as category, color, pattern, and demographic. Monolithic embeddings mix these signals into a single vector, making attribute-specific control difficult at retrieval time. Many existing semantic-ID methods provide discrete item codes, but these codes are typically optimized as item-level or residual addresses and do not expose named, independently controllable attribute slots. We introduce MM-slotgate, a multimodal slot encoder that factorizes Fashion-CLIP text and image embeddings into four named attribute slots. Each slot learns its own text-image gate, so visually grounded attributes such as color and pattern can rely more on image evidence, while taxonomy-oriented attributes such as category and demographic can remain more text-driven. On H&M, using a combined slot-similarity and slot-logit retrieval score, MM-slotgate achieves 0.7566 macro ConstraintSatisfied@10, outperforming equal-weight multimodal fusion (0.7142) and fCLIP text-only retrieval (0.4755). The largest gain is on color, which improves from 0.321 to 0.889 (+0.568 absolute), as the learned color gate assigns 57.4% weight to image evidence. The learned gates are interpretable without modality supervision: color is image-leaning, category is text-leaning, and pattern and demographic lie near the middle. The resulting slots also remain controllable: linear probes show no measured excess leakage beyond the label-correlation baseline, and quantized slot codes support targeted intervention, including a 15.3x lift for color. These results suggest that controllable fashion retrieval benefits from typed, attribute-conditioned multimodal slots rather than either a single global embedding or opaque item-level semantic IDs.