发表机构
Walmart Global Tech; University of Illinois at Urbana-Champaign(沃尔玛全球技术公司; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对电商搜索页面级布局决策的离线评估难题,探究语言模型作为可扩展评估器,发现基于表示的方法在预测用户参与度上优于基于提示的方法,为相关评估提供了可靠基础。
AI 中文摘要
电商搜索页面是数百万在线购物者的关键接触点。传统搜索引擎会返回排序后的结果列表,而现代电商搜索页面越来越多地融入推荐系统模块,例如在特定位置展示替代产品分组的次级堆叠。次级堆叠若放置得当,可提升用户参与度;但放置不当可能会打乱浏览流程并降低主结果质量。与已建立了交织等评估技术的传统搜索排序不同,评估页面级布局变化(例如何时何地插入次级堆叠)在不进行成本高昂的在线A/B测试的情况下颇具挑战性。为解决此问题,我们研究离线方法以评估给定布局决策(具体为在特定位置是否加入次级堆叠)是否对用户有益。我们将语言模型作为可扩展评估器进行研究,比较了基于直接提示、提示衍生特征及基于表示的方法。结果显示,在预测用户参与度方面,基于表示的方法始终优于基于提示的判断,表明它们为电商搜索中的离线布局评估提供了可靠基础。
英文摘要
E-commerce search pages are critical touchpoints for millions of online shoppers. While traditional search engines return a ranked list of results, modern E-commerce search pages increasingly incorporate recommender system modules -- for example, secondary stacks that surface alternative product groupings at specific positions. When introduced appropriately, secondary stacks can improve user engagement; however, suboptimal placement may disrupt browsing flow and degrade the primary results. Unlike traditional search ranking, where evaluation techniques such as interleaving are well established, evaluating page-level layout changes e.g., when and where to insert a secondary stack remains challenging without costly online A/B testing. To address this, we study offline methods for evaluating whether a given layout decision -- specifically, the inclusion of a secondary stack at a particular position -- is beneficial to users. We investigate language models as scalable evaluators by comparing direct prompt-based, prompt-derived feature, and representation-based methods. Our results show that representation-based approaches consistently outperform prompt-based judging in predicting user engagement, suggesting they provide a reliable foundation for offline layout evaluation in E-commerce search.
CommentsAccepted at the OARS Workshop, ACM RecSys 2026