无效度的测量:智能体AI评估中的复合可靠性问题
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
浏览论文内容
中文总结 AI 辅助
该研究指出智能体AI评估存在复合可靠性问题,任务生成、LLM模拟、评分者间信度三个层面的失效相乘降低效度,提出八项心理测量学改进建议以提升评估可信度。
中文摘要 AI 辅助
智能体AI系统通过自动化基准进行评估,其得分用于证明部署决策、安全认证和合规性主张的合理性。我们的实证分析表明,这些得分的系统性可信度低于当前实践所认可的水平。该问题存在三个复合层面:第一,任务越来越多地由语言模型生成:对10个流行基准的审计发现,其中7个存在效度缺陷,所有10个均存在报告缺口;第二,人类用户被大语言模型(LLM)模拟器取代,但校准研究记录了模拟器间的差异高达9个百分点,且存在系统性的方向性校准偏差,尤其针对非标准美式英语使用者;第三,我们对55篇论文的结构化调查发现,约82%的论文应用了结构不匹配、不完整或缺失的评分者间信度(IRR)指标。这些失效是相乘而非相加复合的。在独立性假设下,一个在任务生成阶段保留70%有效信号、模拟阶段保留80%、判断阶段保留65%的流程,针对目标构念的效度最多为36%;该界限在实证估计范围内为0.22至0.54。我们将此形式化为$V_{\text{total}} \times V_1 \times V_2 \times V_3$,并证明当同一模型系列作用于所有三个层面时,在相关失效下该界限会进一步收紧。我们基于心理测量学科学提出八项建议:模拟校准下限为组内相关系数(ICC(A,1))≥0.70;按后果级别划分的领域分层信度阈值(α≥0.67/0.70/0.80);基于流程设计的结构化IRR指标选择规则;以及将IRR作为强制报告字段。测量工具已存在,该领域的任务是应用它们。
英文摘要
Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. We present a three-layer compounding validity model, $V_{total} \leq V_1 \times V_2 \times V_3$, that captures multiplicative degradation across task generation ($V_1$), human-simulator calibration ($V_2$), and automated judgment ($V_3$). Under empirically grounded estimates, a pipeline retaining 70% validity at each stage is at most 34% valid against the intended construct (range 0.17-0.54 across the empirical estimate bounds). We examine the model's predictions against a structured survey of 55 published agentic evaluation papers, finding that approximately 82% of papers in this purposive sample apply structurally mismatched, incomplete, or absent inter-rater reliability (IRR) metrics, a pattern consistent with systematic $V_3$ collapse. We further identify empirical evidence of $V_1$ failures (task validity flaws in 7 of 10 popular benchmarks) and $V_2$ miscalibration (up to 9 percentage points inter-simulator variance, with systematic demographic disparities for non-Standard American English speakers). We derive eight prescriptions grounded in psychometric science and domain-stratified reliability thresholds (ICC $\geq$ 0.70; $α\geq$ 0.67/0.70/0.80 by consequence level) that practitioners and benchmark authors can apply immediately. The framework provides a tractable knowledge-based tool for diagnosing and correcting evaluation pipeline validity before deployment decisions are made.
发表机构
- Red Hat, Inc.(红帽公司)
- Alma Mater Europaea University(欧洲母校大学)
机构由 AI 辅助整理,请以论文原文为准。