arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29803cs.IRcs.AI

SEEK:面向工业搜索的具有可进化知识的技能路由评估

SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search

Zhongxin Huang, Songyang Li, Renzhe Zhou, Feiran Zhu, Chenglei Dai, Zhen Xiao, Xuanping Li, Jingwei Zhuo

首次发表
浏览论文内容

中文总结 AI 辅助

针对工业搜索评估标准多维且动态演化的问题,提出SEEK框架,通过技能库外部化标准、动态路由和两阶段训练实现页面级评估与归因,在快手4亿日活场景显著提升评估质量。

中文摘要 AI 辅助

搜索质量评估为工业搜索系统的开发和迭代提供了必要的监督和诊断信号。尽管大语言模型(LLMs)为人工评估提供了一种可扩展的替代方案,但可靠的自动评估仍然具有挑战性:用户在页面级别体验搜索结果,而适用的评估标准是多维且不断演化的。将所有评估标准打包到一个统一的提示中会引入无关上下文和潜在的标准干扰,而通过后训练将这些标准内化则会将规则更新与昂贵的模型重训练周期紧密耦合。为解决这些问题,我们提出了具有可进化知识的技能路由评估(SEEK)。具体而言,SEEK将特定的搜索评估标准外部化到技能库中,为每个查询-结果列表对动态路由相关技能,并采用任务自适应的列表式评估器来生成页面级判断和失败模式归因。一个两阶段的训练流程教会评估器将评估标准与人类偏好对齐,而一个重放门控技能库允许在不进行模型重训练的情况下纳入重复出现的评估知识缺口。在工业短视频搜索上的实验表明,SEEK提高了列表式质量评估的准确性,并在归因诊断方面取得了显著进展。SEEK已在快手(一个拥有超过4亿日活跃用户的短视频平台)部署,显著提升了在线搜索评估的规模和质量。

英文摘要

Search quality evaluation provides essential supervision and diagnostic signals for the development and iteration of industrial search systems. Although large language models (LLMs) offer a scalable alternative to manual assessment, reliable automatic evaluation remains challenging: users experience search results at the page level, while the applicable evaluation criteria are multi-dimensional and continuously evolving. Packing all evaluation criteria into a unified prompt introduces irrelevant context and potential criterion interference, whereas internalizing them through post-training tightly couples rule updates with costly model retraining cycles. To address these issues, we propose Skill-routed Evaluation with Evolvable Knowledge (SEEK). Specifically, SEEK externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution. A two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining. Experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis. SEEK has been deployed at Kuaishou, a short-video platform with over 400 million daily active users, significantly improving the scale and quality of online search evaluation.

发表机构

  • Peking University(北京大学)
  • Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

↑