VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
机构 * Walmart Global Tech(沃尔玛全球技术)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments Accepted at RecSys 2025; DOI:https://doi.org/10.1145/3705328.3748064