arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RetrievalFormer:用于高效近似最近邻检索和冷物品推荐的双编码器Transformer

Keeping the Index Open: The Recommendation-Side Cost of Shared Search and Recommendation

Theodore Rogers, Joe Standerfer, Dmitrii Timoshenko, Haoxue Li, Zuhaib Akhtar, Soyoung Yang

arXiv 2608.24079首次发表:更新:

发表机构

Amazon Web Services(亚马逊网络服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RetrievalFormer双编码器Transformer,在高效近似最近邻检索框架下,实现对冷物品的推荐,其冷启动推荐效果优于多数基线方法,但热态推荐性能弱于部分重新训练基线。

AI 中文摘要

共享搜索与推荐索引必须仅从特征对新物品打分,因为搜索没有探索位。在覆盖同一目录下两种场景的公开日志中,38.6%的保留查询-搜索展示包含从未展示或访问过的物品。对于用户冷交互,基于特征的塔(tower)满足了这一需求,与99个采样负样本相比无明显损失(Recall@20为0.9595,而热态为0.9510)。词汇基线达到相似的对等水平,而全目录检查在统计上仍不确定。因此,双编码器检索使索引对新物品保持“开放”,这与需要重新训练的ID-softmax推荐器不同。我们在推荐中对这种“开放性”定价,针对六个顺序基线,每个基线通过五轮修正目标进行重新训练和调优。一个float32时间戳错误打乱了19.7%用户的留一法目标。在MovieLens-1M上,热态准确率比最强的重新训练基线低5.2%的Recall@20和11.4%的NDCG@20;在MIND上,与五个最强基线的差距缩小至0.8%至3.6%,尽管该模型在七个模型中排名第六。在严格的零泄露冷启动评估下,内容塔在无冷特定训练的情况下达到0.172±0.006的Recall@20,是最强的重新训练专用方法(0.124±0.007)的1.4倍,是无训练基准的3倍。精确的全softmax训练在MIND-small上使Recall@20提升54%,在MovieLens-1M上比采样InfoNCE提升6.9%,但每步需重新计算全目录,且在240K物品时耗尽加速器内存。近似最近邻检索无法解释剩余差距,服务成本未比ID-softmax检索下降,而历史窗口扫描可解释配方后剩余差距的一半。目录规模下的精确质量训练仍是未解决的问题。

英文摘要

A shared search-and-recommendation index must score new items from features alone because search has no exploration slot. In a public log covering both surfaces over one catalog, $38.6\%$ of held-out query-search impressions show an item never previously shown or visited. For user-cold engagements, the feature-based tower serves this demand without measurable loss against $99$ sampled negatives ($0.9595$ Recall@20 versus $0.9510$ warm). A lexical baseline reaches similar parity, while a full-catalog check remains statistically undecided. Dual-encoder retrieval therefore keeps the index \emph{open} to new items, unlike an ID-softmax recommender that requires retraining. We price this openness on recommendation against six sequential baselines, each retrained and tuned through five rounds on corrected targets. A float32 timestamp bug had reordered leave-one-out targets for $19.7\%$ of users. On MovieLens-1M, warm accuracy trails the strongest retrained baseline by $5.2\%$ Recall@20 and $11.4\%$ NDCG@20. On MIND, the gap narrows to $0.8$--$3.6\%$ relative to the five strongest baselines, though the model ranks sixth of seven. Under strict zero-leakage cold-start evaluation, the content tower achieves $0.172 \pm 0.006$ Recall@20, $1.4\times$ the strongest retrained dedicated method ($0.124 \pm 0.007$) and $3\times$ a training-free floor, without cold-specific training. Exact full-softmax training raises Recall@20 by $54\%$ on MIND-small and $6.9\%$ on MovieLens-1M over sampled InfoNCE, but recomputes the full catalog each step and exhausts accelerator memory at $240$K items. Approximate nearest-neighbor search explains none of the remaining gap, serving cost does not regress against ID-softmax retrieval, and a history-window sweep explains half the post-recipe remainder. Exact-quality training at catalog scale remains the open problem.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑