arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01833cs.AIcs.SE

持续过程级评估:面向演进中的企业AI智能体技能

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin

首次发表
浏览论文内容

中文总结 AI 辅助

针对企业AI智能体技能演进中的过程级行为漂移,提出结合结果级与过程级检查的持续评估框架,通过模板测试和LLM评判器检测偏差,实验显示多数通过最终检查的运行仍存在过程偏差,依赖归因可有效减少根因。

中文摘要 AI 辅助

企业AI智能体技能会随着工具API、模型和规范的变化而演进,然而仅基于最终输出的评估可能遗漏过程级的行为漂移。我们提出了一种结合结果级与过程级检查的持续评估框架,并将其应用于企业价值感知弹性系统中一个业务价值判定技能的收入与生产力变体。该框架独立计算每次运行的真实值,物化可复用的模板测试,并通过程序化检查及一个范围受限的LLM评判器评估工具选择、参数、执行顺序和数据库完整性。我们在两个技能、两个规范变体、两个智能体框架和三个模型上评估了240次试验。在通过所有适用的最终数值检查的175次试验中,有162次(92.6%;Wilson 95%置信区间:87.7-95.6%)包含另一个评估器检测到的偏差。在更宽泛的七项最终状态定义下,164次通过运行中有151次(92.1%;95%置信区间:86.9-95.3%)仍违反了轨迹检查。依赖归因将每次运行的平均失败检查数从6.34个减少到2.65个根因。规范敏感性因模型和框架而异,探索性自助法交互区间在全部三次收入比较和三次生产力比较中的一次中排除了零。运行时解析的模板在评估配置中提供了可复用的回归覆盖;在真实API演进下的纵向验证仍是未来工作。

英文摘要

Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.

发表机构

  • IBM
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑