arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

最近邻但非最亲:面向纯内容搜索与推荐共享策展人反馈基础设施

Nearest but Not Dearest: Shared Curator-Feedback Infrastructure for Content-Only Search and Recommendation

Matt Sandler

arXiv 2609.30568首次发表:更新:

AI 中文总结

针对纯内容音乐发现平台,利用策展人反馈将拒绝分为声音与上下文失败,分别路由至嵌入重加权头和约束过滤器,将拒绝率从38.17%降至28.83%,证明共享反馈基础设施可同时改进搜索与推荐。

AI 中文摘要

一个已部署的B2B音乐发现平台,在单一授权目录、单一LAION-CLAP联合音频-文本嵌入空间、单一候选生成过滤器和单一排序头上,同时服务于查询驱动的搜索(文本提示、氛围标签)和种子驱动的推荐(种子曲目和艺术家电台),且两条路径均不消费终端听众的行为信号。在这种纯内容机制下,策展人判断是可用的主要反馈信号,而离线余弦相似度对其预测效果不佳:38%的余弦最近邻被策展人拒绝。拒绝情况揭示了一个清晰的分区:大多数(55%)是编码器可以解决的声音失败(风格、节奏、情绪不匹配),而相当多的少数(37%)是与波形正交的上下文失败(错误语言、节日内容、宗教内容、权利和歌词标志)。我们将这种声音与上下文分解部署为反馈基础设施,将每种失败模式路由到能够吸收它的层:上下文失败路由到候选生成时的约束过滤器,声音失败路由到表示层的嵌入重加权头——两者均位于搜索/推荐分叉之下,因此单个策展人循环即可维护两种体验。在相隔一个月两次生产轮次中收集的1,200条策展人判断上,联合干预将拒绝率从38.17%降至28.83%(相对下降24.5%,McNemar卡方=22.4,p=2.2e-6)。会计分解将4.08个百分点的下降归因于过滤器合格类别,5.25个百分点归因于其余部分;部署是非盲且复合的,因此这是生产会计界限,而非因果估计。我们将其作为工业案例研究而非验证过的一般方法呈现,并以纯内容发现的教训作结:失败分区与范式分区正交,修正落在两个范式共享的层中。

英文摘要

A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generation filter, and one ranking head -- and neither path consumes end-listener behavioral signal. In this content-only regime, curator judgment is the principal feedback signal available, and offline cosine similarity predicts it poorly: 38% of cosine-nearest neighbors are rejected by curators. The rejections reveal a clean partition: a majority (55%) are sound failures the encoder could address (style, tempo, mood mismatch), and a substantial minority (37%) are context failures orthogonal to the waveform (wrong language, holiday content, devotional content, rights and lyric flags). We deploy this sound-vs-context decomposition as feedback infrastructure, routing each failure mode to the layer that can absorb it: context failures to a constraint filter at candidate generation, sound failures to an embedding reweighting head at the representation layer -- both below the search/recommendation split, so a single curator loop maintains both experiences. On 1,200 curator judgments collected over two production rounds one month apart, the combined intervention reduces rejection rate from 38.17% to 28.83% (-24.5% relative, McNemar chi-squared = 22.4, p = 2.2e-6). An accounting decomposition attributes 4.08 pp of the drop to filter-eligible categories and 5.25 pp to the rest; the deployment was unblinded and compound, so this is a production accounting bound, not a causal estimate. We present this as an industrial case study rather than a validated general method, and close with lessons for content-only discovery: the failure partition is orthogonal to the paradigm partition, and corrections land in layers shared by both paradigms.

Comments8 pages, 3 figures, 3 tables. Accepted for oral presentation at the Unified Search and Recommendation Workshop (USRW) at RecSys 2026; workshop is non-archival

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑