arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24467cs.AI

HMGCLIP:面向电商表征学习的异构多粒度对比学习

HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning

Qiuyu Zhu, Yi Gao, Zhichao Wan, Mingyang Ma

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有多模态大语言模型难以捕获电商产品细粒度属性的问题,提出异构多粒度对比学习框架HMGCLIP,构建异构超图对齐多粒度语义,发布细粒度电商数据集,在多基准上性能优于对比方法。

中文摘要 AI 辅助

尽管近期多模态大语言模型(MLLMs)在通用产品理解方面取得了进展,但它们会将产品信息隐式编码为全局嵌入,从而限制了其捕获细粒度属性的能力。这一局限会影响需要精确属性区分的任务性能,例如区分视觉相似产品间细微的材质差异。为应对这一挑战,我们提出了统一多模态嵌入框架HMGCLIP。通过构建异构超图,我们利用超图拓扑挖掘具有结构感知的难负样本,并在关系和超边层面对齐多粒度语义。该设计支持双粒度推理机制,可动态融合属性证据以适配细粒度和粗粒度下游任务。此外,我们发布了一套全面的细粒度电商数据集,以助力未来的基准测试。在该新数据集及公开MAVE基准上开展的大量实验表明,HMGCLIP的性能优于强大的多模态编码器、MLLMs及电商基线模型,验证了其优越性。

英文摘要

Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. This limitation hinders performance in tasks requiring precise attribute discrimination, such as distinguishing subtle material differences among visually similar products. To address this challenge, we propose HMGCLIP, a unified multimodal embedding framework. By constructing a heterogeneous hypergraph, we leverage hypergraph topology to mine structure-aware hard negatives and align multi-granular semantics at both relation and hyperedge levels. This design enables a dual-granularity inference mechanism that dynamically fuses attribute evidence for both fine-grained and coarse-grained downstream tasks. Furthermore, we release a comprehensive fine-grained e-commerce dataset to facilitate future benchmarking. Extensive experiments on this new dataset and the public MAVE benchmark show that HMGCLIP outperforms strong multimodal encoders, MLLMs, and e-commerce baselines, validating the superiority of HMGCLIP.

发表机构

  • Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)

机构由 AI 辅助整理,请以论文原文为准。

↑