AI 中文总结
研究生产云技能测试覆盖率问题,开发测量管道恢复操作义务、映射测试用例并生成覆盖报告,结合模型辅助与专家评审确保覆盖率基于证据,将覆盖差距转化为测试改进建议。
AI 中文摘要
云平台越来越多地提供可重复使用的云技能,以指导人工智能代理创建、配置、监控和排除云资源故障。面向工作流的技能指定了多个资源操作、用户选择、验证步骤和恢复行为。现有技能评估主要衡量技能是否提高任务成功率。然而,通过现有测试用例并不能揭示技能所指定的哪些行为从未经过测试。本文引入了技能测试覆盖率,用于衡量技能测试套件对其操作测试义务的覆盖程度。我们定义了覆盖单元、测试用例覆盖关系、计算过程和声明边界。由于这些单元和链接在自然语言技能包中不明确,我们开发了一个测量管道,用于恢复操作义务、组织其工作流上下文、使用用户提示和初始资源状态映射测试用例,并生成可审计的覆盖报告。模型辅助候选生成与专家评审相结合,以确保报告的覆盖率基于原始技能和测试用例证据。然后,一个评审后模块将确认的覆盖差距转化为基于源的测试改进建议。
英文摘要
Cloud platforms increasingly deliver reusable Cloud Skills that guide AI agents through multi-step resource operations, user choices, validation, and recovery. Existing Skill evaluation primarily measures whether a Skill improves task success, but passing the available testcases does not reveal which behaviors specified by the Skill remain untested. We introduce Skill Test Adequacy, a scenario-conditioned criterion that evaluates a test suite against the complete set of operational test obligations specified by a Skill. Given a Skill package and normalized testcases containing a prompt, an initial resource state, and expected user decisions, the assessment determines whether each obligation is exercised by at least one testcase scenario; the resulting records provide both a suite-level score and explicit test gaps. We operationalize the criterion through parallel obligation proposals, disagreement-preserving aggregation, testcase-level status proposals, expert review, and source-grounded recommendations. Alibaba Cloud deploys this process as a mandatory gate before task-success evaluation and subsequent release checks. Among 157 initial assessments recorded before gate-driven remediation, 57 (36.3%) fall below the mandatory 80% gate and 76 (48.4%) remain below the recommended 90% level. The process also produces 132 reports containing 639 obligation-level recommendations, with a median of four per Skill. Finally, we release SkillAdeqBench, an exploratory subset of the reviewed records for studying automatic adequacy assessment. Skill Test Adequacy complements task-success evaluation by making the untested scope of production Cloud Skills explicit.