arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从有效到有用:递归自改进推荐的验证后获取

From Valid to Useful: Post-Verification Acquisition for Recursive Self-Improving Recommendation

Tonmoy Hasan, Taylor Foust, Shao Tang, Leonardo Neves, Aman Gupta, Hiroto Udagawa, Helder Dias, Daniel Silva, Rohan Ramanath

arXiv 2610.04302首次发表:更新:

发表机构

Nubank(Nubank)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对递归自改进推荐中验证后序列选择问题,提出分歧感知的DA-RSIR方法,利用BALD分数排序并限制源序列贡献,在24项比较中全面优于保留全部策略,确立验证后获取为独立控制点。

AI 中文摘要

序列推荐器可以生成合成交互序列,并在增强语料库上进行重训练,形成一个递归自改进循环。为限制错误累积,现有方法会验证每个生成序列是否仍能预测用户的真实交互,并丢弃偏离真实交互的序列。然而,验证并不决定哪些已验证序列应训练下一个模型。当每个已验证序列都用于训练时,产生更多已验证序列或更长延续的源序列具有更大影响力,尽管这两个数量都不表明这些序列对下一个模型的帮助程度。我们将决定哪些已验证序列用于训练下一个模型的问题表述为“验证后获取”,并引入{\f 分歧感知的递归自改进推荐(DA-RSIR)}。DA-RSIR限制每个源序列的贡献,并根据模型对其增强交互的预测分歧程度对已验证序列进行排序。它使用基于贝叶斯主动学习分歧(BALD)的分数,并通过蒙特卡洛(MC)dropout进行估计。DA-RSIR不需要额外标签、教师模型或质量评分器。在四个数据集、三个推荐模型和两个指标上,它在全部24个比较中优于保留全部方法,并在24个总体比较中的23个中取得最高均值;两个指标上的总体改进均具有统计显著性。单轮DA-RSIR超过了保留全部方法在五轮递归中的最佳增益。这些发现将验证后获取确立为递归自改进中的一个独立控制点,将哪些序列通过验证与哪些已验证序列用于训练下一个模型分开。

英文摘要

Sequential recommenders can generate synthetic interaction sequences and retrain on the augmented corpus in a recursive self-improvement loop. To limit error accumulation, current methods verify each generated sequence remains predictive of the user's real interactions and discard those that drift away from it. Verification does not, however, determine which verified sequences should train the next model. With every verified sequence used for training, source sequences yielding more verified sequences or longer continuations have more influence, although neither quantity indicates how much those sequences will help the next model. We formulate the decision of which verified sequences are used to train the next model as \emph{post-verification acquisition} and introduce {\bf Disagreement-Aware Recursive Self-Improving Recommendation (DA-RSIR)}. DA-RSIR caps each source sequence's contribution and ranks its verified sequences by how much the model's predictions disagree over their augmented interactions. It uses a score derived from Bayesian Active Learning by Disagreement (BALD) and estimated with Monte Carlo (MC) dropout. DA-RSIR requires no extra labels, teacher model, or quality scorer. Across four datasets, three recommender models, and two metrics, it improves on the retain-all approach in all $24$ comparisons and attains the highest mean in $23$ of $24$ overall; the aggregate improvement is statistically significant on both metrics. A single DA-RSIR round exceeds the retain-all approach's best gain over five recursive rounds. These findings establish post-verification acquisition as a separate control point in recursive self-improvement, separating which sequences pass verification from which verified sequences are used to train the next model.

Comments15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑