发表机构
Target Corporation(塔吉特公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种基于LinUCB的上下文多臂老虎机系统,可实时优化电商商品页面布局,在零售平台A/B测试中较启发式基线实现会话级性能正向提升,为Web学习型设计提供可扩展路径。
AI 中文摘要
电商平台正通过机器学习日益个性化用户体验,但页面布局决策仍以静态规则和人工策划为主。我们提出一种可扩展的基于多臂老虎机(bandit)的系统,该系统能实时优化商品页面布局,同时保留人类对设计意图的控制权。上下文多臂老虎机模型会利用用户、商品及类别级别的特征,为每个会话动态选择最有效的布局。该系统采用基于LinUCB的策略,在从实时用户交互中学习的过程中平衡探索与利用。其架构专为无缝集成到大规模Web服务栈而设计,支持低延迟推理和持续模型更新。该系统首先在入门级商品页面上进行测试,在某大型零售平台的在线A/B部署中,我们的方法在会话级性能指标上较强大的启发式基线取得了正向提升。我们的结果表明,上下文多臂老虎机可有效优化商品发现的视觉和结构方面以提升用户参与度,为实现Web的学习型设计提供了可扩展的路径。
英文摘要
E-commerce platforms increasingly personalize user experiences through machine learning, yet page layout decisions remain dominated by static rules and manual curation. We present a scalable bandit-based system that optimizes product page layouts in real time while preserving human control over design intent. A contextual bandit model dynamically selects the most effective layout for each session using user, item, and category-level features. The system leverages a LinUCB-based policy to balance exploration and exploitation as it learns from live user interactions. The architecture is designed for seamless integration into large-scale web serving stacks, supporting low-latency inference and continuous model updates. The system was first tested on entry product pages. In online A/B deployments on a major retail platform, our approach achieved positive lifts in session-level performance metrics over a strong heuristic baseline. Our results demonstrate that contextual bandits can effectively optimize visual and structural aspects of product discovery for user engagement, providing a scalable path toward learning-to-design the web.
CommentsAccepted to The Web Conference 2026 (short paper track), but later withdrawn due to internal prioritization. Subsequently accepted to the Online & Adaptive Recommender Systems Workshop (held in conjunction with the 20th ACM Conference on Recommender Systems, RecSys 2026)