arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30333cs.IRcs.AI

超越排序准确率:针对下一个购物篮回购推荐的大语言模型引用特征理由评估

Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation

Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对下一个购物篮回购推荐,探究LLM作为独立推荐器的可行性,发现其排序能力不及监督排序器,但可作为经验证的解释组件,且理由质量需与排序准确率分开评估。

中文摘要 AI 辅助

下一个购物篮回购推荐通常被表述为一项排序任务:给定顾客的购买历史,系统对可能再次被需要的已购商品进行排序。然而在生产环境中,排序准确率仅是推荐质量的组成部分之一,顾客也可能从关于当前推荐某商品的简洁证据中获益。大语言模型(LLM)提供了一种可行途径,可通过基于特征的、可被人类理解的理由来呈现此类证据,这些理由以可解释的行为信号为基础。我们构建了涵盖购买周期、频率、时效性、用户行为及商品流行度的回购特征,并在两个公开的杂货数据集和一个私有零售数据集上评估LLM。我们研究两点:一是与启发式排序器和监督排序器相比,现成的LLM是否可作为下一个购物篮评分器使用;二是LLM引用的特征是否具备基于结果的排序信号。对于后者,我们采用跨模型特征掩码协议,在掩码选定特征后测量排序退化情况,将LLM引用的特征与特定模型的归因方法进行比较。结果显示,LLM的评分无法与监督排序器相媲美,表明现成的LLM不应被用作独立的回购推荐器。不过,即便排序性能未提升,调整提示词和证据表示仍可在部分场景下改善基于结果的特征掩码结果;该效果依赖于数据集,且未与归因基线始终匹配。这些发现表明,LLM可作为经验证的解释组件发挥实际作用,而非主要排序器,其理由质量需与排序准确率分开评估。

英文摘要

Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one component of recommendation quality. Customers may also benefit from concise evidence about why an item is recommended now. Large language models (LLMs) offer a potential way to surface such evidence through feature-based, human-readable rationales grounded in interpretable behavioral signals. We construct repurchase features spanning cadence, frequency, recency, user behavior, and item popularity, and evaluate LLMs on two public grocery datasets and one proprietary retail dataset. We investigate (1) whether off-the-shelf LLMs can use these features as next-basket scorers relative to heuristic and supervised rankers, and (2) whether LLM-cited features carry outcome-grounded ranking signal. For the latter, we compare LLM-cited features with model-specific attribution methods under a cross-model feature-masking protocol that measures ranking degradation after masking selected features. Our results show that LLM scores are not competitive with supervised rankers, suggesting that off-the-shelf LLMs should not be used as standalone repurchase recommenders. However, changes in prompt and evidence representation can improve outcome-grounded feature-masking results in some settings even when ranking performance does not improve; the effect is dataset-dependent and does not consistently match attribution baselines. These findings suggest a practical role for LLMs as validated explanation components rather than primary rankers, with rationale quality evaluated separately from ranking accuracy.

发表机构

  • Walmart Global Tech(沃尔玛全球技术)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑