arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MetaStrategy:基于可执行大语言模型策略的生成式排序

MetaStrategy: Generative Ranking with Executable LLM Strategies

Chengyu Lai, Jiuning Lin, Zhibo Xiao, Xiaodong Zhu, Ruiquan Lan, Bin Zhang, Zihong Huang, Wendong Zhang, Chuxin Chen, Yinjiang Cai, Shuai Zhong, Lingqing Zhang, Dimin Wang, Jialin Zhu, Han Zhu

arXiv 2608.09440首次发表:更新:

AI 中文总结

MetaStrategy是一种可执行LLM策略的生成式排序框架,经淘宝部署测试,其提升了多项核心业务指标且未增加响应时间。

AI 中文摘要

工业推荐系统需在用户、业务、商业及体验目标耦合的场景下对异构内容进行排序。现有生成式排序方法通常直接构建物品序列,难以与成熟预测模型、运营规则及领域级防护措施集成。本文提出MetaStrategy框架,该框架生成结构化的可执行排序策略:基于请求上下文,大语言模型(LLM)策略输出类型化JSON包,控制目标权重、内容与类别偏好、体验约束及位置策略;确定性验证器与编译器实例化一个隔离的生成器,该生成器与现有模型在生成器-评估器(GE)架构的列表级评估器下原子竞争。我们在生产路径重放环境中训练该策略,该环境通过当前重排序栈重新执行已记录的请求,且不会暴露用户。该方法结合选择、相对排序及基线提升奖励,将频繁策略作为竞争者反馈的自竞争课程,以及由评估器路由的奖励增强在线策略蒸馏,将互补的4B参数教师模型迁移至紧凑的0.8B参数学生模型。我们通过差异触发的近线生成将MetaStrategy部署至淘宝首页“猜你喜欢”功能,LLM推理保持在同步排序之外,响应时间(RT)无明显增加。在为期七天的用户随机在线A/B测试中,MetaStrategy在处理侧GE调用中获胜27.93%,并显著提升点击量(click PV)2.11%、商品详情页浏览量(IPV)3.12%及交易额2.83%。

英文摘要

Industrial recommender systems rank heterogeneous content under coupled user, business, commercial, and experience objectives. Existing generative ranking methods typically construct item sequences directly, making them difficult to integrate with mature predictive models, operational rules, and field-level guardrails. We present MetaStrategy, a framework that instead generates a structured, executable ranking strategy. Conditioned on request context, a large language model (LLM) policy emits a typed JSON bundle controlling objective weights, content and category preferences, experience constraints, and position policies. A deterministic validator and compiler instantiate an isolated Generator that competes atomically with incumbents under the list-level Evaluator of the Generator-Evaluator (GE) architecture. We train the policy in a production-path replay environment that re-executes logged requests through the current re-ranking stack without user exposure. The method combines selection, relative-rank, and baseline-lift rewards, a self-competitive curriculum that feeds frequent strategies back as competitors, and Evaluator-routed reward-augmented on-policy distillation that transfers complementary 4B-parameter Teachers into a compact 0.8B-parameter Student. We deploy MetaStrategy in Taobao Homepage Guess You Like through diff-triggered nearline generation; LLM inference remains outside synchronous ranking, with no observable increase in response time (RT). In a seven-day user-randomized online A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and significantly improves click page views (click PV) by 2.11%, item-detail page views (IPV) by 3.12%, and transaction amount by 2.83%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑