AI 中文总结
研究基于网易云音乐部署的大语言模型驱动音乐推荐代理Melo,针对实体幻觉和长尾退化故障模式,采用推理时实体基础和反思性重试机制,经测试在播放列表留存和参与度指标上有提升,强调运行时机制对推荐进展的重要性。
AI 中文摘要
我们描述了Melo,一个部署在网易云音乐上的由大语言模型驱动的音乐推荐代理。Melo被构建为一个基于异构工具的确定性五节点状态图,具有由提示和状态机驱动的编排策略,而非经过微调的控制器。在工业规模下,瓶颈不在于大脑有多智能,而在于系统如何检测并从大脑所犯的错误中恢复。两种生产故障模式推动了设计:实体幻觉,即代理采用未得到实时目录或用户行为索引支持的解释;以及长尾退化,即过度受限的请求退化为通用的流行回退。我们用两种互补机制解决它们。推理时实体基础将生产搜索索引重新用作验证原语,在实体决策向下游传播之前对其进行把关。反思性重试将工具链故障的原因用语言表达出来,并将其输入到下一个规划步骤,以便系统能够放宽或修改约束,而不是盲目回退。在网易云音乐播放列表界面进行的为期一个月的在线A/B测试报告称,主要播放列表留存指标提升了超过2个百分点,核心播放列表参与度指标提升了超过一分钟。离线消融在我们的评估集上从三层基础堆栈中分离出实体错误识别率降低7.8个百分点,对我们评估集的触发会话分析显示,反思性重试在5.8%的会话中触发,流程级恢复率为59%。我们的部署经验表明,在这种规模下由大语言模型驱动的音乐推荐进展,在很大程度上既取决于能够捕捉并纠正大脑错误的命名可消融运行时机制,也取决于大脑本身:这是一个我们供社区测试的假设。
英文摘要
We describe Melo, an LLM-powered music recommendation agent deployed on NetEase Cloud Music. Melo is structured as a deterministic five-node state graph over heterogeneous tools, with a prompt- and state-machine-driven orchestration policy rather than a fine-tuned controller. At industrial scale, the bottleneck is not how smart the brain is but how the system detects and recovers from the mistakes that brain makes. Two production failure modes drove the design: entity hallucination, where the agent commits to interpretations unsupported by the live catalog or user-behavior index, and long-tail degradation, where over-constrained requests collapse to generic popular fallbacks. We address them with two complementary mechanisms. Inference-time entity grounding repurposes the production search index as a verification primitive that gates entity decisions before they propagate downstream. Reflective retry verbalizes failure reasons from a broken tool chain and feeds them into the next planning step, so the system can relax or revise constraints rather than fall back blindly. A one-month online A/B test across NetEase Cloud Music's playlist surfaces reports an over 2 pp lift in a primary playlist retention metric and a lift of over one minute in a core playlist engagement metric. Offline ablation isolates a 7.8 pp reduction in entity misidentification from the three-layer grounding stack on our evaluation set, and a triggered-session analysis on our evaluation set shows reflective retry firing on 5.8% of sessions with 59% process-level recovery. Our deployment experience suggests that progress on LLM-powered music recommendation at this scale depends as much on the named, ablatable runtime machinery that catches and corrects the brain's mistakes as on the brain itself: a hypothesis we offer for the community to test.
Journal refProceedings of the 20th ACM Conference on Recommender Systems (RecSys '26), September 27-October 02, 2026, Minneapolis, MN, USA