AI 中文总结
针对电子商务搜索排名中多模态信息利用的局限,提出MMRM框架,通过共享主干从多信号学习生成多重商品表示,引入多重用户表示策略,并经实验验证其有效性,已成功应用于京东搜索引擎提升性能。
AI 中文摘要
多模态信息对电子商务搜索排名至关重要。现有工作通常通过协作信号微调通用多模态大语言模型(MLLMs)来利用多模态数据,然后将派生表示作为商品特征集成到排名模型中。但这些方法存在两个主要局限:依赖单一协作信号微调MLLM,未利用多任务排名所需的异构信号;将多模态表示视为常规商品特征,未充分利用其对用户行为建模的潜在潜力。为应对这些挑战,我们提出多重多模态表示模型(MMRM),这是一个将MLLMs与多样协作信号对齐的统一框架。通过使用带有特定任务令牌和投影层的共享主干,MMRM能从多个信号同时学习并在单次推理中生成全面的多重商品表示。此外,我们在排名模型中引入多重用户表示策略,通过利用多重商品表示的基于搜索的行为序列建模来派生特定任务的用户表示。大量实验证明了MMRM的卓越效率和有效性。值得注意的是,MMRM已成功部署在京东电子商务搜索引擎中,为数百万日常用户带来显著性能提升。
英文摘要
Multimodal information is pivotal for e-commerce search ranking. Existing works leverage multimodal data typically by fine-tuning general Multimodal Large Language Models (MLLMs) via collaborative signals, subsequently integrating the derived representations into ranking models as item features. Despite their efficacy, these methods face two primary limitations: (1) they rely on a single collaborative signal for MLLM fine-tuning, failing to exploit the heterogeneous signals essential for multitask ranking; and (2) they treat multimodal representations as regular item features in ranking models, underutilizing their latent potential for user behavior modeling. To address these challenges, we propose the Multiplex Multimodal Representation Model (MMRM), a unified framework that aligns MLLMs with diverse collaborative signals. By employing a shared backbone with task-specific tokens and projection layers, MMRM simultaneously learns from multiple signals and generates comprehensive multiplex item representations in a single inference pass. Furthermore, we introduce a multiplex user representation strategy in ranking models, which derives task-specific user representations via search-based behavior sequence modeling leveraging multiplex item representations. Extensive experiments demonstrate MMRM's superior efficiency and effectiveness. Notably, MMRM has been successfully deployed in the JD e-commerce search engine, yielding significant performance gains for millions of daily users.
CommentsAccepted by SIGIR2026