arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sona技术报告

Sona Technical Report

Sona Team, Alexandr Udeneev, Aleksei Krasilnikov, Alexey Nadtochiy, Andrey Semenov, Andrey Tsyrkunov, Anna Krivonos, Anna Lipkina, Artem Matveev, Daniil Burlakov, Daniil Leshchev, Daria Tikhonovich, Denis Burshtein, Ekaterina Dmitrieva, Eugene Krofto, Grigorii Khlystov, Ilya Murzin, Kirill Golovko, Ksenia Sycheva, Leonid Dmitriev, Mariia Rozaeva, Mariia Ulianova, Mikhail Sandul, Nikolai Savushkin, Oleg Sorokin, Roman Odobesku, Semyon Panenko, Sergei Liamaev, Sergei Makeev, Vadim Shilov, Veronika Ivanova, Viktor Yanush, Vladimir Baikalov, Vladislav Dodonov, Vladislav Tytskiy

arXiv 2608.11015首次发表:更新:

AI 中文总结

本文提出单模型生成式推荐系统Sona,在A/B测试中替代Yandex Music原有多阶段推荐级联,显著提升活跃用户、收听时长等核心指标,效果优于此前最强模型Argus。

AI 中文摘要

我们推出Sona,这是一款用于Yandex Music的单模型生成式推荐系统。在在线A/B测试中,Sona替换了整个生产级联系统,该系统包含超过15个候选生成器,以及后续的预排序和排序模型,这些模型会使用数百个特征,包括来自大型Transformer模型(如Argus)和目标注意力评分器的信号,同时显著提升了关键互动指标。Sona的架构围绕共享用户表示统一了候选生成和排序过程。其编码器将用户按时间顺序排列的已记录互动事件序列转换为隐藏状态,供自回归解码器和排序模块使用。下一个token预测和蒸馏目标共同更新编码器,通过相同的用户状态将生成与排序耦合。Sona及其教师排序器均未使用手动设计的特征;两者均基于已记录的事件字段和学习到的物品表示运行。在最终的Sona配置中,更大的教师模型在训练期间提供排序目标,但在服务阶段不参与,仅部署编码器、解码器和排序模块作为单个模型。我们在针对智能扬声器上My Vibe(Yandex Music最大的推荐界面之一)的实时流量开展的在线A/B实验中评估了Sona。与生产对照组相比,Sona在主要指标活跃用户数上实现了4.53%的统计显著提升,总收听时长提升了6.30%,点赞数提升了11.42%。这些效果是在之前部署带来的改进基础上的增量提升。活跃用户数的提升幅度是Sona部署前该界面最强模型Argus此前所带来增量的2.35倍。这些结果表明,联合训练的单模型可以替代成熟的多阶段推荐级联系统,同时在实时流量上提升推荐质量。

英文摘要

We introduce Sona, a single-model generative recommender for Yandex Music. In an online A/B test, Sona replaced the entire production cascade, comprising more than 15 candidate generators followed by pre-ranking and ranking models that consume hundreds of features, including signals from large transformer models such as Argus and target-attention scorers, while significantly improving key engagement metrics. The architecture of Sona unifies candidate generation and ranking around a shared user representation. Its encoder transforms the user's chronological sequence of logged engagement events into hidden states consumed by both the autoregressive decoder and the Ranking Module. The next-token-prediction and distillation objectives jointly update the encoder, coupling generation and ranking through the same user state. Neither Sona nor its Teacher Ranker uses hand-engineered features; both operate on logged event fields and learned item representations. In the final Sona configuration, the larger teacher supplies ranking targets during training but is absent from serving, leaving the encoder, decoder, and Ranking Module as a single deployed model. We evaluate Sona in an online A/B experiment using live traffic from My Vibe on smart speakers, one of Yandex Music's largest recommendation surfaces. Relative to the production control, Sona produced statistically significant uplifts of 4.53% in Active Users, the primary metric, 6.30% in Total Listening Time, and 11.42% in Likes. These effects were incremental to improvements retained from preceding deployments. The Active Users uplift was 2.35 times the increment previously delivered by Argus, the strongest model deployed on this surface before Sona. These results show that a single jointly trained model can replace a mature multi-stage recommendation cascade while improving recommendation quality on live traffic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑