arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪些应该进入评估集?面向智能体可扩展性平台的、基于能力分类法的回归评估集编建流水线

Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

Tezan Sahu, Aritra Das, Pankaj Mittal, Sudipta Das

arXiv 2608.01004首次发表:更新:

发表机构

Microsoft; Indian Institute of Technology Roorkee(微软; 鲁尔基印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体可扩展性平台的回归评估集编建问题,提出基于能力分类法的流水线,通过三个组件实现查询决策,可适配带类型化能力分类法的各类编建场景。

AI 中文摘要

托管智能体可扩展性界面的平台团队面临回归经济学悖论:每个新入驻客户都会提供针对其领域优化的评估集,但平台的回归评估集必须严格控制在与发布节奏挂钩的查询数量上限内。据我们所知,目前已发表的工业级流水线均未解决这一平台侧编建问题:现有评估框架均为客户侧框架,而基准压缩研究将基准视为固定集合而非不断流入的评估集流。我们描述了一个应用于Microsoft 365 Copilot中带自定义动作的声明式智能体的、基于能力分类法的编建流水线。该流水线以智能体规范和客户评估集为输入,将每个查询映射到平台拥有的能力分类法中,并输出每个查询的决策(接纳、删除、替换或人工审核),其核心思路是:健康的回归评估集应是能捕获最大范围能力签名的最小查询集——能力签名是查询共同使用的不同能力组合。该流水线包含三个组件:一是分类器,通过确定性规范提取与大语言模型(LLM)语义推理相结合的混合方法,生成每个(查询、能力)对的判定;二是调用质量(IQ)评估器,对查询使用每个能力的彻底程度进行评分,因此与现有条目共享签名的新查询仍可被识别为更优测试并将其替换;三是整合器,通过基于规则的决策级联,将流入的查询与回归评估集在覆盖范围和质量上进行比较,背后由保守编建器支撑,该编建器仅建议删除条目。该机制与分类法无关,适用于任何带有类型化能力分类法的回归评估集编建问题,包括那些会因流水线呈现的证据而演变的分类法。

英文摘要

Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count ceiling bounded by release cadence. To our knowledge, no published industrial pipeline addresses this platform-side curation problem: existing evaluation frameworks are customer-side, and benchmark-compression work treats benchmarks as fixed pools rather than streams of incoming sets. We describe a capability-taxonomy-driven curation pipeline applied to declarative agents with custom actions in Microsoft 365 Copilot. It takes an agent specification and a customer's eval set as input, projects each query into a platform-owned capability taxonomy, and outputs per-query decisions (admit, drop, swap, or human review), under the philosophy that a healthy regression set is the minimal set of queries capturing the maximal spread of capability signatures -- distinct combinations of capabilities a query exercises together. Three components instantiate this: a classifier producing per-(query, capability) verdicts via a hybrid of deterministic specification-based extraction and large-language-model (LLM) semantic inference; an Invocation Quality (IQ) rater scoring how thoroughly a query exercises each capability, so a new query sharing a signature with an existing entry can still be recognized as a better test and displace it; and a consolidator comparing incoming queries against the regression set on coverage and quality through a rule-based decision cascade, backed by a conservative curator that only suggests evictions. The mechanism is taxonomy-agnostic and applies to any regression eval-set curation problem with a typed capability taxonomy, including taxonomies that evolve in response to the very evidence the pipeline surfaces.

CommentsExtended abstract accepted at the SERI 2026 Industry Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑