arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当智能体指标衡量不同事物时:对Praxa AI流水线的基于证据的审计

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline

Stefan G. Creadore, Peyton Woakz

arXiv 2609.12017首次发表:更新:

AI 中文总结

通过审计Praxa AI流水线,揭示数值正确但构念错位的指标问题,提供源关联案例与可复用验证包,区分门控、时序与核算声明。

AI 中文摘要

智能体评估可能在数值上正确,但衡量的构念与其标签所暗示的构念不同。我们对选定的Praxa AI实现文件、历史评估工件和操作记录进行了回顾性测量审计。一份139例离线路由报告包含112次通过和27次失败,尽管门控失败为零,因为已知缺口被明确豁免于门控。一份无标识符导出代表了8,843行工具尝试记录:8,395个记录时长和448个缺失值。在这些时长中,121个等于有符号32位最大值并带有弃用客户端标签;检查的数据库代码将生命周期时长钳制到该最大值。合并记录的第99百分位为2,147,483,647毫秒,而服务器观察到的已完成调用的第99百分位为38,118.31毫秒。这是分层对比,而非处理效应。在一个记录在案的单轨迹压缩试点中,报告的后续输入减少率为94.39%,但触发和后续调用合计的减少率为46.54%。我们复现了描述性计算,通过独立的加权有理算术实现验证了91个时序统计量,并执行了13个评分函数和12个分析验证器测试。有限完成边界表明,缺失时长如何限制全行时序陈述,而无需插补值。本研究的贡献在于提供一个源关联的案例研究和可复用的验证包,用于将门控策略、生命周期时序和请求级核算与更广泛的智能体性能声明区分开来。历史提供方运行和完整当前流水线未被独立复现;一般能力优越性和总体统计显著性未被确立。

英文摘要

Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records. A 139-case offline routing report contains 112 passes and 27 failures despite zero gating failures, because known gaps are explicitly exempted from the gate. An identifier-free export represents 8,843 tool-attempt rows: 8,395 recorded durations and 448 missing values. Of the durations, 121 equal the signed 32-bit maximum and carry abandoned-client labels; inspected database code clamps elapsed lifecycle age. The pooled recorded 99th percentile is 2,147,483,647 ms, versus 38,118.31 ms among server-observed completed calls. This is a stratum contrast, not a treatment effect. In a documented single-trajectory compaction pilot, the reported follow-up input reduction is 94.39%, but the reduction across the trigger and follow-up calls together is 46.54%. We reproduce the descriptive calculations, verify 91 timing statistics through a separate weighted rational-arithmetic implementation, and execute 13 scoring-function and 12 analysis-verifier tests. Finite-completion bounds show how missing durations limit all-row timing statements without imputing values. The contribution is a source-linked case study and reusable verification package for separating gate policy, lifecycle timing, and request-level accounting from broader agent-performance claims. Historical provider runs and the full current pipeline were not independently reproduced; general capability superiority and population-level statistical significance are not established.

Comments22 pages, 4 figures. Retrospective measurement audit with aggregate data, offline analysis code, and a prospective, unexecuted operator-coverage catalog in ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑