arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

审计序列推荐中的语义增益:一种轻量级恢复测试

Auditing Semantic Gains in Sequential Recommendation: A Lightweight Recovery Test

Kong Wang, Zhongke He, Xiang Chen, Hongwei Zeng, Kai Deng, Long Wang, Kehua Yang

arXiv 2608.01260首次发表:更新:

发表机构

Hunan University; Dalian University of Technology; Tongji University; University of Chinese Academy of Sciences; Xinjiang College of Science & Technology; Beihang University(湖南大学; 大连理工大学; 同济大学; 中国科学院大学; 新疆科技学院; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出轻量级可审计恢复测试LIME-Rec,结合三类专家模型在三个Amazon数据集上验证语义推荐增益源于真实物品-文本对应,为序列推荐的语义增益归因提供了透明方法。

AI 中文摘要

近期,基于语义和生成式检索的推荐系统相较于仅使用ID的序列基线取得了显著提升,但目前仍不清楚这些增益是源于语言模型推理、语义-ID生成、端到端语义架构、更强的离线物品表示,还是源于语义与协同信号的互补。我们通过LIME-Rec(一种轻量级且可审计的恢复测试)探究这种归因模糊性。LIME-Rec结合了三个独立专家:SASRec序列专家、ItemCF共现专家,以及基于冻结的BAAI/bge-base-en-v1.5物品嵌入的语义专家。它们的全目录分数按用户归一化后,通过可审计的分数级融合,再经有界历史校准进行组合。融合门和校准头仅在验证数据上拟合,无需服务时的语言模型推理,且各专家的贡献可单独检查。在Amazon Beauty、Toys和Sports数据集上,LIME-Rec的R@10分数分别为0.0996、0.1105和0.0593,优于最强对比基线7.0%-12.0%。未进行历史校准的三专家融合始终优于校准后的SASRec,表明仅校准无法解释这种恢复。在物品ID间随机置换物品文本嵌入会使R@10降低13.6%-17.5%,说明增益依赖于真实的物品-文本对应关系,而非额外的表示能力。这些结果表明,在将改进归因于服务时语言建模、语义-ID生成或更重的语义机制之前,应排除从离线物品表示进行轻量级恢复和透明融合的可能。

英文摘要

Recent semantic and generative-retrieval recommenders report substantial improvements over ID-only sequential baselines, but it remains unclear whether these gains arise from language-model reasoning, semantic-ID generation, end-to-end semantic architectures, stronger offline item representations, or complementary semantic and collaborative signals. We investigate this attribution ambiguity through LIME-Rec, a lightweight and auditable recovery test. LIME-Rec combines three independent experts: a SASRec sequential expert, an ItemCF co-occurrence expert, and a semantic expert based on frozen BAAI/bge-base-en-v1.5 item embeddings. Their full-catalog scores are normalized per user and combined through auditable score-level fusion followed by bounded history calibration. The fusion gate and calibration head are fitted on validation data only, require no serving-time language-model inference, and keep each expert contribution separately inspectable. On Amazon Beauty, Toys, and Sports, LIME-Rec achieves R@10 scores of 0.0996, 0.1105, and 0.0593, outperforming the strongest comparison baseline by 7.0%-12.0%. Three-expert fusion without history calibration consistently outperforms calibrated SASRec, showing that calibration alone does not explain the recovery. Randomly permuting item-text embeddings across item IDs reduces R@10 by 13.6%-17.5%, indicating that the gains depend on genuine item-text correspondence rather than additional representation capacity. These results suggest that lightweight recovery from offline item representations and transparent fusion should be ruled out before improvements are attributed to serving-time language modeling, semantic-ID generation, or heavier semantic machinery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑