arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当历史误导时:指令引导的LLM生成式推荐的不对称边际监督

When History Misleads: Asymmetric Margin Supervision for Instruction-Guided LLM Generative Recommendation

Ming Yin, Yuhan Yang, Chen Chen, Xinyu Lin, Wentao Shi, Fangcong Yin, Chaofei Yang, Chao Yang, Jiyan Yang, Hui Zhang, Ning Jiang, Yiran Chen, Qifan Wang

arXiv 2610.02600首次发表:更新:

发表机构

Duke University; Meta(杜克大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对指令引导生成式推荐中历史事件误导当前请求的问题,提出AIMS方法,通过冻结参考模型生成边际监督并采用不对称损失,在六个LLM骨干上提升Recall和NDCG。

AI 中文摘要

在指令引导的生成式推荐中,基于LLM的推荐器需要平衡两个目标:响应用户当前请求并与其交互历史中的偏好保持一致。当两者冲突时,历史事件可能覆盖请求。我们表明,将单个历史事件的影响转化为监督面临两个障碍。首先,对推荐影响最大的事件不一定支持目标物品。其次,移除误导性事件可以提高目标的得分,但竞争物品的得分可能提高更多,因此更高的目标得分本身并不能保证更好的排名。我们提出非对称干预引导边际监督(AIMS),将移除单个历史事件的影响转化为排名监督。对于已正确排名的训练请求,冻结参考模型识别特定于请求的删除,这些删除同时提高目标的得分及其在推荐截止点附近相对于竞争者的边际。这些边际作为训练目标,同时完整历史保留为输入。训练结合交叉熵与不对称辅助损失,该损失惩罚边际不足并仅通过竞争者得分路由其梯度。推理不变,无需历史编辑或删除搜索。在工业数据集和两个公共基准上的六个LLM骨干中,AIMS在Recall和NDCG上优于强基线。消融支持特定于请求的边际和不对称监督,所选删除优先移除违反约束的历史。

英文摘要

In instruction-guided generative recommendation, LLM-based recommenders need to balance two goals: responding to the user's current request and aligning with the preferences in their interaction history. When the two conflict, history events can override the request. We show that turning the effect of individual history events into supervision faces two obstacles. First, the events that most influence a recommendation are not necessarily the ones that support the target item. Second, removing a misleading event can raise the target's score but a competing item's score even more, so a higher target score alone does not guarantee a better ranking. We propose Asymmetric Intervention-Guided Margin Supervision (AIMS), which converts the effect of removing individual history events into ranking supervision. For training requests already ranked correctly, a frozen reference model identifies request-specific deletions that improve both the target's score and its margin over a competitor near the recommendation cutoff. These margins serve as training targets, while the complete history is retained as input. Training combines cross-entropy with an asymmetric auxiliary loss that penalizes margin shortfalls and routes its gradient only through the competitor score. Inference is unchanged, requiring no history editing or deletion search. Across six LLM backbones on an industrial dataset and two public benchmarks, AIMS improves Recall and NDCG over strong baselines. Ablations support request-specific margins and asymmetric supervision, and the selected deletions preferentially remove constraint-violating history.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑