arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10124cs.IRcs.LG

生成式推荐中带遗漏目标的训练:将监督与概率竞争分离

Training with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Xuesi Wang, Yangbin Shi, Xiaolin Zheng

AI总结:

本研究针对生成式推荐中训练附加遗漏目标与推理候选竞争的问题,提出分离监督与概率竞争的匹配损失,实验表明消除竞争可提升FT-NDCG,建议按生成器评估候选补全。

AI中文摘要:

生成式推荐器返回有限的候选集,并可能在重排序前遗漏已观察到的目标。一种训练策略将这些遗漏目标附加到重排序训练列表中,尽管推理时仍仅对原始候选进行排序。此操作同时改变了检索目标的权重,增加了对附加目标的监督,并使两组目标在概率上产生竞争。因此,附加与不附加的对比无法解释返回项排序的变化。我们构建了三种匹配的损失函数,在保持检索目标权重固定的同时,分别引入附加目标监督和组间竞争。中间损失在两组内进行训练但分别归一化,防止仅训练用目标与推理候选竞争。使用已发布的OneRec模型和本地训练的Amazon生成器进行的实验表明,这种竞争可能损害返回项的排序。在四个预设的Amazon Video Games比较中,消除这种竞争使全目标归一化折损累计增益(FT-NDCG)提升了7.8%至22.2%;在用户层面95%的置信区间以及四次训练运行中三次的区间均排除了零。一个保守的开发集规则在保留的一个类别中为三个生成器中的两个选择了附加目标训练,并在另一个类别中对全部三个生成器予以拒绝,避免了1.7%的损失。因此,候选补全应针对每个生成器进行评估,而非自动应用。

英文摘要:

Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2\%; 95\% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7\% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.

补充信息

↑