arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生产云技能是否得到充分测试?在实践中衡量和管理技能测试覆盖率

Are Production Cloud Skills Adequately Tested? Measuring and Governing Skill Test Adequacy in Practice

Haotian Si, Junyi Chen, Shuyang Yu, Ruifeng Nie, Jiate Li, Jianqiang Zhao, Meng Li, Dengcheng He

arXiv 2607.22015首次发表:更新:

AI 中文总结

研究生产云技能测试覆盖率问题,开发测量管道恢复操作义务、映射测试用例并生成覆盖报告,结合模型辅助与专家评审确保覆盖率基于证据,将覆盖差距转化为测试改进建议。

AI 中文摘要

云平台越来越多地提供可重复使用的云技能,以指导人工智能代理创建、配置、监控和排除云资源故障。面向工作流的技能指定了多个资源操作、用户选择、验证步骤和恢复行为。现有技能评估主要衡量技能是否提高任务成功率。然而,通过现有测试用例并不能揭示技能所指定的哪些行为从未经过测试。本文引入了技能测试覆盖率,用于衡量技能测试套件对其操作测试义务的覆盖程度。我们定义了覆盖单元、测试用例覆盖关系、计算过程和声明边界。由于这些单元和链接在自然语言技能包中不明确,我们开发了一个测量管道,用于恢复操作义务、组织其工作流上下文、使用用户提示和初始资源状态映射测试用例,并生成可审计的覆盖报告。模型辅助候选生成与专家评审相结合,以确保报告的覆盖率基于原始技能和测试用例证据。然后,一个评审后模块将确认的覆盖差距转化为基于源的测试改进建议。

英文摘要

Cloud platforms increasingly deliver reusable Cloud Skills that guide AI agents through multi-step resource operations, user choices, validation, and recovery. Existing Skill evaluation primarily measures whether a Skill improves task success, but passing the available testcases does not reveal which behaviors specified by the Skill remain untested. We introduce Skill Test Adequacy, a scenario-conditioned criterion that evaluates a test suite against the complete set of operational test obligations specified by a Skill. Given a Skill package and normalized testcases containing a prompt, an initial resource state, and expected user decisions, the assessment determines whether each obligation is exercised by at least one testcase scenario; the resulting records provide both a suite-level score and explicit test gaps. We operationalize the criterion through parallel obligation proposals, disagreement-preserving aggregation, testcase-level status proposals, expert review, and source-grounded recommendations. Alibaba Cloud deploys this process as a mandatory gate before task-success evaluation and subsequent release checks. Among 157 initial assessments recorded before gate-driven remediation, 57 (36.3%) fall below the mandatory 80% gate and 76 (48.4%) remain below the recommended 90% level. The process also produces 132 reports containing 639 obligation-level recommendations, with a median of four per Skill. Finally, we release SkillAdeqBench, an exploratory subset of the reviewed records for studying automatic adequacy assessment. Skill Test Adequacy complements task-success evaluation by making the untested scope of production Cloud Skills explicit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑