发表机构
Michigan State University; Intuit(密歇根州立大学; Intuit公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对推荐系统中负样本缺乏可解释证据的问题,提出一种基于符号规则和LLM解释的隐式负候选发现方法,在工业与公开数据集上提升精确度与下游性能。
AI 中文摘要
推荐系统从观察到的用户-物品交互中学习,但显式负面反馈通常不可用。由于深度学习模型需要负信号进行训练,负采样方法通常将选定的未观察交互视为负样本。然而,缺失的交互并不能解释用户为何对某个物品不感兴趣,或者是否有足够的证据将其标记为负样本。这在商业推荐中尤为重要,因为负信号应具有可解释性并与业务目标一致。我们提出了隐式负候选发现,以识别由观察到的客户行为支持但未观察到的交互。我们将这些模式编码为符号规则,根据支持度、信息量和产品相关性进行评分,并按证据对保留的规则进行排序。然后,LLM 使用业务目标和领域知识解释保留的规则;解释与统计证据结合在最终报告中。我们在工业 B2B 环境和五个公开推荐数据集上评估了我们的方法。在工业环境和三个公开数据集中的候选质量评估显示,与评估的基线相比,精确度更高,而符号选择在工业任务中,每个正例四个负例的情况下,将下游测试 PR-AUC 比随机选择提高了 12.5%。我们的结果表明,负候选有效性可以与下游推荐性能分开评估。这种区分使得基于证据、业务对齐且可解释的负选择成为可能,在稀疏、偏斜的真实世界推荐设置中提高了可解释性和模型训练。
英文摘要
Recommender systems learn from observed user-item interactions, but explicit negative feedback is often unavailable. Since deep learning models require negative signals for training, negative sampling methods typically treat selected unobserved interactions as negatives. However, a missing interaction does not explain why a user is uninterested in an item or whether there is sufficient evidence to label it negative. This is especially important in business recommendation, where negative signals should be interpretable and aligned with business objectives. We formulate implicit negative candidate discovery to identify unobserved interactions supported by observed customer behavior. We encode these patterns as symbolic rules, score them based on support, informativeness, and product relevance, and rank the retained rules by evidence. An LLM then interprets the retained rules using business objectives and domain knowledge; the interpretations are combined with the statistical evidence in the final report. We evaluate our method in an industrial B2B setting and across five public recommendation datasets. Candidate-quality evaluations in the industrial setting and three public datasets show higher precision than the evaluated baselines, while symbolic selection improves downstream test PR-AUC by 12.5% over random selection with four negatives per positive example in the industrial task. Our results show that negative candidate validity can be evaluated separately from downstream recommendation performance. This distinction enables evidence-based, business-aligned, and explainable negative selection, improving both interpretability and model training in sparse, skewed, real-world recommendation settings.