NEXT:通过视觉语言模型进行推理驱动的视频推荐
NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model
浏览论文内容
中文总结 AI 辅助
研究提出推理驱动的视频推荐框架NEXT,通过对用户已观看视频推理推断其下一个意图来检索后续视频。训练NEXT-8B视觉语言模型,在DocVQA性能上表现出色,在特定评估中提升下一个意图逻辑质量,部署后带来生产收益,证明其可作为实用推理引擎。
中文摘要 AI 辅助
我们提出了NEXT(Next-interest EXploration Transformer),这是一个推理驱动的视频推荐框架,它对用户刚刚观看的视频进行推理,推断观看者的下一个意图,并检索具体的后续视频。明确的延续内容如剧集直接链接;隐含情况通过生成意图查询和搜索匹配候选来处理。这种从项目到意图再到项目的公式产生了超越共同参与相关性或语义相似性的定向推荐。为使该框架在大规模下可靠,我们训练了NEXT-8B,一个经过专门训练的8B视觉语言模型,采用三阶段方法:用于查询无关证据提取的感知增强强化学习、对真实和合成视觉问答混合进行分布对齐的监督微调,以及用于最后阶段对齐的组相对策略优化。NEXT-8B在单模型DocVQA性能上最佳,在特定任务的LLM作为评判的评估中,下一个意图逻辑质量比基础模型提高了3.3%。我们将NEXT作为大规模社交媒体推荐系统中的额外检索路径进行部署,观察到了显著的生产收益。总体而言,NEXT表明经过精心训练的紧凑视觉语言模型可作为生产规模下下一个兴趣探索的实用推理引擎。
英文摘要
We present NEXT (Next-interest EXploration Transformer), a reasoning-driven video recommendation framework that reasons over the video a user has just watched, infers the viewer's next intent, and retrieves concrete follow-up videos. Explicit continuations such as episodes are linked directly; implicit cases are handled by generating intent queries and searching for matching candidates. This Item-to-Intent-to-Item formulation produces directed recommendations beyond co-engagement correlation or semantic similarity. To make this framework reliable at scale, we train NEXT-8B, a purpose-trained 8B vision-language model with a three-stage recipe: Perception-Enhanced Reinforcement Learning for query-agnostic evidence extraction, Distribution-Aligned Supervised Fine-Tuning over real and synthetic visual QA mixtures, and Group Relative Policy Optimization for last-mile alignment. NEXT-8B achieves the best single-model DocVQA performance, ranking second overall only behind a multi-agent system while surpassing a substantially larger 200B+ scale model, and improves next-intent logic-wise quality by 3.3% over the base model in a task-specific LLM-as-a-judge evaluation. We deploy NEXT as an additional retrieval path in a large-scale social media recommendation system and observe statistically significant production gains, including +0.53% watch time and +0.51% distinct video exposure. Overall, NEXT shows that a carefully trained compact vision-language model can serve as a practical reasoning engine for next-interest exploration at production scale.
发表机构
- Meta Platforms, Inc.(元平台公司)
机构由 AI 辅助整理,请以论文原文为准。