arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CALMRec:用于长期推荐的因果对齐语言记忆

CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation

Gengyu Zhan

arXiv 2607.23647首次发表:更新:

发表机构

Shenzhen University(深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对当前大语言模型推荐器的问题,提出CALMRec框架,利用冻结多模态语言模型转换信息,维护多种记忆,通过倾向加权更新等减少偏差,在多环境评估中表现良好,提升了折扣长期价值,释义基准上语义原子NDCG显著提高。

AI 中文摘要

大语言模型(LLMs)可汇总异构用户证据,但当前基于LLM的推荐器常将持久偏好、瞬时意图和曝光诱导行为混为一谈,使推荐易受反馈循环影响。本文提出一种用于长期推荐的模型无关框架。该方法使用冻结的多模态语言模型将商品内容和反馈转换为基于证据的语义原子,维护单独的短期、长期和曝光记忆。倾向加权更新减少策略诱导的曝光偏差,保守离线评论家在行为支持约束下对候选商品重新排序以考虑延迟满意度。解释仅使用有影响力的证据原子并通过反事实删除进行检查。在类似电子商务、新闻和短视频的环境中评估该框架,结果显示该方法在折扣长期价值上分别比最强替代方法提高6.1%、7.6%和6.7%。去除倾向校正或保守支持正则化后价值显著下降。在保留的释义基准上,冻结的指令语言模型的语义原子NDCG比TF-IDF增加一倍多。

英文摘要

Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation vulnerable to feedback loops: repeated exposure is mistaken for preference, immediate clicks dominate delayed satisfaction, and fluent explanations need not reflect the ranking decision. We propose our method, a model-agnostic framework for long-horizon recommendation. Our method uses a frozen multimodal language model to convert item content and feedback into evidence-grounded semantic atoms, then maintains separate short-term, long-term, and exposure memories. Propensity-weighted updates reduce policy-induced exposure bias, while a conservative offline critic reranks candidates for delayed satisfaction under a behavior-support constraint. Explanations use only influential evidence atoms and are checked by counterfactual deletion. We provide an identification result and evaluate the framework in e-commerce-like, news-like, and short-video-like environments. Across ten seeds, our method improves discounted long-term value over the strongest alternative by 6.1%, 7.6%, and 6.7%, respectively. Twenty-seed paired ablations show significant value drops after removing propensity correction (0.739 +/- 0.191) or conservative support regularization (0.523 +/- 0.234). A frozen instruction language model also more than doubles semantic-atom NDCG over TF-IDF on a held-out paraphrase benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑