arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20614cs.AI

评估技能,而非仅评估智能体:智能体驱动的技能持续评估

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh… 展开作者

Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee

AI总结:

该研究提出ACES框架,通过配对实际 trials等方式评估技能附加价值,实验表明其能发现扫描式审核无法观测的信号,开源实现已在NVIDIA SkillEvaluator中提供。

AI中文摘要:

企业智能体程序正从原型阶段迈向生产环境,在此环境中,可复用的技能、工具及工作流包必须通过证据而非文字进行审核。当前的审核机制通常仅扫描这些制品的结构、风格和安全性,但无法回答部署层面的核心问题:该能力包是否能在相同模型、沙箱和评分策略下,帮助实际智能体完成企业任务?我们提出ACES(Agentic Continuous Evaluation of Skills,智能体驱动的技能持续评估),一种基于仓库的框架,用于将技能和产品能力包作为可执行智能体制品进行评估。ACES会运行包含与不包含目标技能的配对实际 trials,将轨迹标准化为智能体轨迹交换格式(Agent Trajectory Interchange Format, ATIF),对六个默认运行时指标进行评分,并报告技能增益(Skill Lift):在固定任务、工具集、工作空间和评分器下,目标技能的附加价值。该协议同样支持产品自有任务套件,用于比较基线、技能、组合、团队技能及插件目标。针对来自内部企业仓库和公开目录的145项真实技能,仅扫描式审核能发现有用的创作问题,但测量的是互补维度(结构与LLM评判员的斯皮尔曼相关系数ρ=0.14)。在64项生产技能中的58项及四个主要工具集产生的947个已评分配对案例中,综合技能增益的均值为0.2134(95%配对案例置信区间[0.1967, 0.2301]);仅结果增益(准确率和目标准确率的平均值)为0.1799。72.8%的配对案例中综合增益为正。最大的流程指标增益出现在技能执行、行为检查和技能效率方面——这些是关于发现、路由、工作流遵循和工具使用的信号,是文档扫描无法观察到的。该方法论的开源实现可在NVIDIA SkillEvaluator中获取。

英文摘要:

Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer the deployment question: does the capability package help a live agent complete enterprise tasks under the same model, sandbox, and grading policy? We present ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts. ACES runs paired live trials with and without a target skill, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime metrics, and reports Skill Lift: the target skill's added value for a fixed task, harness, workspace, and scorer. The same protocol supports product-owned task suites that compare baseline, skill, bundle, team-skill, and plugin targets. On 145 real skills from internal enterprise repositories and public catalogs, scan-only gates surface useful authoring issues but measure complementary facets (structural versus LLM-judge Spearman $ρ= 0.14$). Across 947 scored paired cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift is 0.2134 (95\% paired-case CI [0.1967, 0.2301]); mean outcome-only lift, the average of accuracy and goal accuracy, is 0.1799. Composite lift is positive in 72.8\% of paired cases. The largest process-metric gains appear in skill execution, behavior check, and skill efficiency---signals about discovery, routing, workflow following, and tool use that document scans cannot observe. An open-source implementation of the methodology is available in NVIDIA SkillEvaluator.

补充信息

相关深度报道

↑