发表机构
Central South University(中南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对外卖推荐系统多模态融合难、语义与行为对齐不足的问题,提出三阶段流程GALA,通过生成式RL对齐阶段弥合预训练-微调差距,在淘宝商购部署后提升了订单量与AUC等指标。
AI 中文摘要
现代外卖领域的推荐系统日益利用图像、文本和用户交互历史等多模态信号提升用户体验,但有效融合这些异质模态仍具挑战,既阻碍多模态信号的联合建模,也难以适配不断变化的用户意图。主流两阶段方法中,图像-文本编码器的内容语义预训练与行为驱动排序模型的分离,限制了语义理解与用户行为模式的对齐。为解决这些问题,本文提出GALA,一种三阶段流程,其核心创新在于中间的“生成式RL对齐”阶段,该阶段从用户行为构建多模态预训练数据,并通过基于转换的奖励对其进行优化,有效弥合预训练-微调差距以对齐下游目标。GALA包含三个阶段:第一阶段,对搜索日志中的查询-图像-文本对进行感知行为的三元组预训练,以提前捕捉用户意图和内容偏好;第二阶段为新颖的中间阶段,通过奖励驱动优化(GRPO)优化多模态嵌入,使其动态适配用户行为并弥合预训练-微调差距;最后阶段,通过带混合损失的自适应门控融合多模态与ID嵌入,在长期以ID为主导的训练中保留多模态的贡献。GALA已部署在淘宝商购的生产环境中,服务超过2亿日活跃用户。与最先进(SOTA)方法相比,它实现了+0.12/+0.20 AUC的持续离线提升,以及更优的PCOC指标;大规模在线A/B测试进一步显示订单量提升0.55%,证实了GALA在工业规模下的有效性及其对不同需求模式的鲁棒性。
英文摘要
Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline whose core innovation lies in an intermediate "generative RL alignment" stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. GALA comprises three stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, a novel intermediate stage that refines multimodal embeddings through reward-driven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55 percent increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.
Comments13 pages, 12 figures, 5 tables. Accepted at the 2026 IEEE International Conference on Data Engineering (ICDE 2026), Industry and Applications Track