PEAR:面向工业搜索系统的渐进式基于证据的自动研究
PEAR: Progressive Evidence-Based AutoResearch for Industrial Search Systems
浏览论文内容
中文总结 AI 辅助
针对工业搜索中非平稳流量下证据积累与多保真度评估的挑战,提出PEAR框架,通过证据驱动状态机和置信度门控验证阶梯,在真实系统中显著提升主订单/DAU。
中文摘要 AI 辅助
自动研究通过迭代实验改进系统:智能体提出候选修改方案,对其进行评估,并利用评估结果指导后续探索。将该范式应用于工业搜索面临两大挑战。(1)常见的自动研究方法遵循“保留更优者”规则,即保留得分最高的候选方案用于后续实验。在非平稳流量条件下,瞬时收益可能被误认为持续性改进,从而损害搜索知识的可靠积累。(2)候选修改方案可在多个保真度层级上进行评估,从低成本代理指标到在线验证,这些层级在成本、目标一致性和统计可靠性方面存在差异。现有方法依赖单一信号或特定任务流程,缺乏统一依据来跨层级使用证据以指导搜索。我们提出了渐进式基于证据的自动研究(PEAR),包含两个互补组件。证据驱动的自动研究在预定义目标和干预范围内,为每个策略任务维护一个独立、由假设引导的研究状态。每个状态通过“计划-执行-评估-更新”的转换过程演进,该过程将实验与上下文感知的证据解释和假设修订相连接。置信度门控验证阶梯将评估组织为四个保真度递增的层级:离线回放、影子流量评估、快速在线评估和决策级在线评估。统一的基于置信度的门控仅在证据支持统计显著的正向效应时提升候选方案,从而实现广泛的低成本探索,同时将昂贵的在线实验保留给被提升的候选方案。在真实的工业搜索系统中,使用PEAR优化的策略在两个A/B实验中,相对于各自基线,显著提升了主订单/日活跃用户数(Main Order/DAU)2.7336%和3.2957%。
英文摘要
AutoResearch improves systems through iterative experimentation: agents propose candidate modifications, evaluate them, and use the results to guide subsequent exploration. Applying this paradigm to industrial search presents two challenges. (1) Common AutoResearch approaches follow a keep-if-better rule, retaining the highest-scoring candidate for subsequent experiments. Under non-stationary traffic, transient gains may be mistaken for persistent improvements, impairing reliable accumulation of search knowledge. (2) Candidate modifications can be evaluated at multiple fidelity levels, from low-cost proxies to online validation, differing in cost, objective alignment, and statistical reliability. Existing methods rely on individual signals or task-specific procedures, lacking a unified basis for using evidence across levels to guide search. We introduce Progressive Evidence-Based AutoResearch (PEAR) with two complementary components. Evidence-driven AutoResearch maintains an independent, hypothesis-guided research state for each strategy task within a predefined objective and intervention scope. Each state evolves through a Plan-Execute-Evaluate-Update transition that links experimentation to context-aware evidence interpretation and hypothesis revision. Confidence-Gated Verifier Ladder organizes evaluation into four levels of increasing fidelity: Offline Replay, Shadow-Traffic Evaluation, Rapid Online Evaluation, and Decision-Grade Online Evaluation. A unified confidence-based gate promotes candidates only when evidence supports a statistically significant positive effect, enabling broad low-cost exploration while reserving costly online experiments for promoted candidates. In a real-world industrial search system, strategies optimized with PEAR significantly increased Main Order/DAU by 2.7336% and 3.2957% relative to their respective baselines in two A/B experiments.
发表机构
- Global E-Commerce Agentic Search Team(全球电商智能体搜索团队)
机构由 AI 辅助整理,请以论文原文为准。