arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于库存的策略级优化:无需训练的AI搜索

Inventory-Grounded Policy-Level Optimization for Training-Free AI Search

Wei Zhou, Tiandeng Wu, Jiandong Ding, Zhufeng Fan, Yi Cao

arXiv 2609.04813首次发表:更新:

发表机构

Huawei Technologies Co., Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI搜索系统库存频繁更新导致传统方法失效的问题,提出无需训练的IGPO方法,将策略与环境事实分离,经14天A/B测试实现CTR提升3.17%、不良案例减少38.9%。

AI 中文摘要

在部署初期,AI搜索系统通常在频繁更新的产品目录上运行,因此可用商品及其属性不能被视为可编码到固定提示或策略中的稳定知识。微调、强化学习和静态提示补丁表现不佳:标签稀缺,奖励随库存变化而漂移,模型发布成本高昂,提示修复很快失效。我们提出Inventory-Grounded Policy-Level Optimization(IGPO,基于库存的策略级优化),一种适用于固定AI搜索流水线的无需训练的方法。IGPO将策略与环境事实分离:它学习用于处理运行时库存证据的策略指南,而非记忆可用商品。在线阶段,IGPO通过探测库存并构建库存画像来为每个查询提供依据,然后将相关策略指南注入检索和选择提示中。离线阶段,随机滚动按查询分组——混合结果组直接产生对比信号,库存引导的探索循环可区分未命中的检索路径与在观察到的库存证据下无匹配支持的情况。自2026年5月起,IGPO已部署在商业智能助手AI搜索系统中。为期14天的完整IGPO方案在线A/B测试显示,相对点击率(CTR)提升3.17%,审计不良案例减少38.9%。

英文摘要

Early in deployment, an AI search system typically operates over a frequently updated product catalog, so the available items and their properties cannot be treated as stable knowledge that can be encoded in fixed prompts or strategies. Fine-tuning, reinforcement learning, and static prompt patches fit poorly: labels are scarce, rewards drift with inventory, model releases are costly, and prompt fixes quickly stale. We present Inventory-Grounded Policy-Level Optimization (IGPO), a training-free approach for fixed AI search pipelines. IGPO separates policy from environment facts: it learns Policy Guidelines for acting on runtime inventory evidence rather than memorizing available items. Online, IGPO grounds each query by probing the inventory and constructing an inventory portrait, then injects relevant Policy Guidelines into the retrieval and selection prompts. Offline, stochastic rollouts are grouped by query -- mixed outcome groups directly yield contrastive signal, and an inventory-guided exploration loop distinguishes missed retrieval routes from cases where no matching support is found under the observed inventory evidence. Since May 2026, IGPO has been deployed in a commercial smart-assistant AI search system. A 14-day online A/B test of the complete IGPO treatment shows a 3.17% relative CTR lift and a 38.9% reduction in audited bad cases.

CommentsAccepted at the EMNLP 2026 Industry Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑