发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM推荐方法扁平化处理异质行为导致的性能问题,提出MARI模型,通过构建结构化决策记忆库提升推荐性能与可解释性,在相关任务上优于基线方法。
AI 中文摘要
尽管大型语言模型(LLMs)已被应用于推荐系统,但主流方法大多仅建模单一类型的行为(如浏览或购买)。即便整合多种行为,现有方法也会将异质行为扁平化处理为同质的 token 序列,忽略了它们各自的决策角色。这种扁平化无法捕捉复杂决策中的语义层次与语境细微差别,例如价格与质量之间的权衡,导致在涉及高度相似物品的关键“困难选择”场景中性能下降。为弥合这一差距,我们提出 MARI(可解释性增强的记忆型推荐),其预测基于明确的结构化决策证据。MARI 维护一个决策记忆库(DMB),将用户过去的决策依据存档为结构化决策记忆(SDMs):即目标、约束与权衡的简洁记录。这些 SDMs 通过事后决策提炼从异质行为和用户生成内容中离线生成。通过检索相关 SDMs 增强 LLM 推理,MARI 兼具可解释性与可扩展性,且无需处理长原始序列的高昂成本。大量实验表明,MARI 在标准下一个物品预测任务和新引入的困难选择预测任务上均显著优于最先进的基线方法,且通过将记忆构建与在线推理解耦,带来的延迟开销较低。定性分析揭示了用户决策中可操作的、人类可读的见解,标志着迈向推理感知推荐系统的具体一步。
英文摘要
Despite the adoption of large language models (LLMs) in recommendation systems, prevailing approaches mostly model single-type behaviors (e.g., views or purchases). Even when incorporating multiple behaviors, existing methods flatten heterogeneous actions into homogeneous token sequences, ignoring their distinct decision-making roles. This flattening fails to capture semantic hierarchies and contextual nuances in complex decision-making, such as trade-offs between price and quality. Consequently, performance degrades in critical ``difficult-choice'' scenarios involving highly similar items. To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence. MARI maintains a Decision Memory Bank (DMB) that archives users' past rationales as Structured Decision Memories (SDMs): concise records of goals, constraints, and trade-offs. These SDMs are generated offline via Post-Hoc Decision Distillation from heterogeneous behaviors and user-generated content. By retrieving relevant SDMs to augment LLM reasoning, MARI achieves interpretability and scalability without the prohibitive cost of processing long raw sequences. Extensive experiments show MARI significantly outperforms state-of-the-art baselines on standard next-item prediction and a newly introduced Difficult Choice Prediction task, incurring low latency overhead by decoupling memory construction from online inference. Qualitative analyses reveal actionable, human-readable insights into user decision-making, marking a concrete step toward reasoning-aware recommendation systems.