AEVAL:从轶事性到确定性的智能体技能工作流测试
AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows
浏览论文内容
中文总结 AI 辅助
研究针对智能体技能工作流测试缺乏确定性和可重复性的问题,提出AEVAL框架,通过执行器与评分器分离等方法,实现确定性、可重复的测试,给出分层修复建议,能将虚假通过率转换为可重复失败信号并记录修复过程。
中文摘要 AI 辅助
现代智能体系统越来越依赖技能,即教大型语言模型智能体执行领域任务的自然语言和代码可安装包。随着技能库增长,开发者需要每次变更的自动化质量信号,但当前评估多是轶事性的,缺乏可重复性和可比性。我们提出AEVAL,一个集成持续集成(CI)的框架,用确定性、可重复的测试管道取代现有做法。每个技能变更触发测试事件,技能在自动化执行器中根据开发者声明的评估契约运行,发出结构化、有证据支持的质量信号供下游CI处理。关键在于执行器和评分器的结构分离,防止智能体在执行中自我纠正并将修补后输出判定为通过这种微妙但普遍的失败模式。我们的贡献包括:(i)具有每个技能契约和每次运行工件模式的确定性、变更触发评估协议;(ii)将自我纠正偏差形式化为朴素智能体评估器的一种独特失败模式;(iii)执行器/评分器分离及首次尝试评分规则和明确的自我纠正跟踪;(iv)作为内联合并请求注释发布的分层、基于证据的修复建议方案(LV1因果,LV2质量)。在多个智能体软件开发工具包的生产智能体堆栈中的实际技能上进行验证,AEVAL将虚假的100%通过率转换为可重复的首次尝试失败信号,并带有每个执行器修复的可审计记录。
英文摘要
Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task. As skill repositories grow, developers need automated quality signals on every change, yet evaluation today is largely anecdotal: a developer asks an agent to "try the skill," watches a demo, and forms a subjective impression. This yields neither reproducibility across runs nor comparability across versions, and scales poorly to marketplaces where one regression can silently break dozens of downstream workflows. We present AEVAL (Agentic Evaluation), a CI-integrated framework that replaces this practice with a deterministic, reproducible test pipeline for agentic skills. Every skill change triggers a test event: the skill runs against a developer declared evaluation contract inside an automated executor, emitting a structured, evidence grounded quality signal that downstream CI can route on. A key ingredient is a structural separation between executor and grader, preventing a subtle but pervasive failure mode: an agent that silently self-corrects during execution and then grades its own patched outputs as passing. Our contributions are: (i) a deterministic, change-triggered evaluation protocol with per-skill contracts and per-run artifact schemas; (ii) a formalization of self correction bias as a distinct failure mode of naive agentic evaluators; (iii) an executor/grader separation with a first-attempt grading rule and explicit self-correction tracking; and (iv) a tiered, grounded evidence fix suggestion scheme (LV1 causal, LV2 quality) posted as inline merge-request comments. Validated on real skills in a production agentic stack across multiple agent SDKs, AEVAL converts spurious 100% pass rates into reproducible first-attempt fail signals with an auditable record of every executor fix.
发表机构
- nvidia(NVIDIA公司)
机构由 AI 辅助整理,请以论文原文为准。