AI 中文总结
研究针对工业搜索等系统中特征工程复杂的问题,提出提示生成(PG)框架,通过两个JSON文件解耦特征处理逻辑与模型架构,经组织特征和组件实现三个层面加速,并在淘宝搜索应用取得收益,成为迭代框架。
AI 中文摘要
生成式检索已成为工业搜索、推荐和广告系统越来越常用的范式,能带来显著的在线收益。现有工作多将用户行为序列与大语言模型结合来建模用户偏好。但实践中,特征工程对模型有效性仍至关重要,其复杂性减缓了离线迭代,使在线部署繁重且难以复用。为打破特征处理逻辑与模型架构的紧密耦合,我们提出提示生成(PG),这是一个高级分词器和配置驱动框架,通过两个声明性JSON文件将特征处理逻辑与模型架构解耦。PG在四个类型下组织特征,有三个可组合处理组件来组装和压缩异构特征,在三个层面实现加速:快速训练迭代、快速部署、快速在线推理。PG已部署在淘宝搜索上,交易计数在线A/B提升0.47%,GMV提升0.51%,并已作为生成式检索的迭代框架应用于多个淘宝搜索和推荐团队。
英文摘要
Generative retrieval has become an increasingly adopted paradigm for industrial search, recommendation, and advertising systems, delivering significant online gains. Most existing work combines user behavior sequences with large language models (LLMs) to model user preferences. In practice, feature engineering remains critical to model effectiveness, yet its complexity slows offline iteration and makes online deployment heavy and hard to reuse, all under tight online latency budgets. The root cause is a tight coupling between feature-processing logic and model architecture, where every feature change touches the training and serving code and resists reuse across scenarios. To break this coupling, we present Prompt Generation (PG), a high-level tokenizer and configuration-driven framework that decouples feature-processing logic from model architecture through two declarative JSON files, which serve as the single source of truth for both offline training and online serving, ensuring feature consistency across the two stages. Organizing features under four types with three composable processing components to assemble and compress heterogeneous features, PG delivers acceleration at three levels: (1)fast training iteration: feature experiments require only configuration changes, with built-in token compression for ultra-long sequences; (2)fast deployment: a new scenario only needs to conform to the PG schema and plug into a universal pipeline, with no scenario-specific engineering; (3)fast online inference: engine applies unified optimizations over the standardized configuration, reducing PG's overhead to a negligible level. PG has been deployed on Taobao Search with statistically significant online A/B uplifts of +0.47% in transaction count and +0.51% in GMV, and has been applied across multiple Taobao search and recommendation teams as the iteration framework for generative retrieval.