arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更优描述性推理轨迹质量与推荐效果之间的脱节

The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness

Gustavo Penha, Juan Elenter, Claudia Hauff, Hugues Bouchard, Paul Bennett, Mounia Lalmas

arXiv 2608.23154首次发表:更新:

AI 中文总结

本研究通过2×2因子实验对比语义ID与自然语言标题的推理轨迹质量,发现提升描述性推理轨迹质量无法持续改善传统离线推荐效果。

AI 中文摘要

近期研究聚焦于改进生成式推荐的显式自然语言描述性推理轨迹,这包括利用思维链推理增强语义ID(SID)预测的系统。然而,由于SID是不透明的学习标识符而非自然语言,在大语言模型(LLM)可对其进行推理前,需要代价高昂的对齐操作。这提供了一个可控实验环境,其中项目表示(标题 vs. SID)和语义基础(最小 vs. 广泛的SID对齐)均可独立变化。因此,我们在三个亚马逊产品领域开展了一项2×2因子实验,采用共享的Qwen3-1.7B主干模型,首次对语义ID和自然语言标题的描述性推理轨迹质量进行了可控比较。我们发现,在标准监督微调(SFT)和强化学习(RL)训练下,引入显式描述性推理轨迹会降低传统离线推荐效果,尽管自然语言标题能产生更具语义基础和可解释性的轨迹。广泛的SID对齐可提升描述性轨迹质量,但无法提升传统离线推荐效果,而更丰富的奖励信号可部分恢复性能。总体而言,我们的结果表明,在本研究采用的训练目标和评估协议下,仅提升描述性推理轨迹质量本身不足以持续改善传统离线推荐效果。

英文摘要

Recent work has focused on improving explicit natural-language descriptive reasoning traces for generative recommendation. This includes systems that augment semantic ID (SID) prediction with chain-of-thought reasoning. However, because SIDs are opaque learned identifiers rather than natural language, they require costly alignment before an LLM can reason over them. This provides a controlled experimental setting in which both item representation (Title vs. SID) and semantic grounding (minimal vs. extensive SID alignment) can be varied independently. We therefore present the first controlled comparison of descriptive reasoning trace quality across semantic IDs and natural-language titles in a 2 x 2 factorial study on three Amazon product domains using a shared Qwen3-1.7B backbone. We find that introducing explicit descriptive reasoning traces reduces traditional offline recommendation effectiveness under standard SFT and RL training, even though natural language titles produce substantially more grounded and interpretable traces. Extensive SID alignment improves descriptive trace quality but not traditional offline recommendation effectiveness, while a richer reward signal partially recovers performance. Overall, our results show that improving descriptive reasoning trace quality is not, by itself, sufficient to consistently improve traditional offline recommendation effectiveness under the training objectives and evaluation protocols studied here.

CommentsAccepted at the Recsys'26 Workshop on Agentic and Generative AI for E-Commerce

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑