发表机构
Texas A&M University–Corpus Christi; University of California, Riverside; University of Missouri(德克萨斯A&M大学科珀斯克里斯蒂分校; 加利福尼亚大学河滨分校; 密苏里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文指出将合规监控数据误作比较评估基准是AI评估中的测量有效性问题,并以自动驾驶为例提出评估契约以明确假设。
AI 中文摘要
我们认为,在已部署AI系统的评估中,一个反复出现的失败发生在当为运营监控或监管合规而收集的数据被解释为是为比较评估而设计的时候。自动驾驶提供了一个具体的例子来说明这个问题。美国的脱离和碰撞报告制度产生了有价值的运营证据,但报告范围、暴露程度、部署领域、事件捕获和比较器构建的差异限制了仅从这些测量中能够支持的安全声明。我们将这个问题框定为AI评估中的测量有效性问题,而不是交通特定的数据限制。我们认为,关于已部署AI系统的比较声明要求预期能力、测量结果、暴露机会、部署领域、数据生成过程和评估比较器之间保持一致。以自动驾驶安全评估为案例研究,我们提出了一个评估契约,在运营数据被解释为比较性能的证据之前,使这些假设明确化。更广泛的影响是,用于监控已部署AI系统的数据并不自动成为评估它们的有效基准。
英文摘要
We argue that a recurring failure in the evaluation of deployed AI systems occurs when data collected for operational monitoring or regulatory compliance are interpreted as if they were designed for comparative evaluation. Automated driving provides a concrete example of this problem. U.S. disengagement and crash-reporting regimes produce valuable operational evidence, but differences in reporting scope, exposure, deployment domain, event capture, and comparator construction limit the safety claims that can be supported from these measurements alone. We frame this issue as a measurement-validity problem in AI evaluation rather than as a transportation-specific data limitation. We argue that comparative claims about deployed AI systems require alignment between the intended capability, measured outcome, exposure opportunity, deployment domain, data-generation process, and evaluation comparator. Using automated-driving safety evaluation as a case study, we propose an evaluation contract that makes these assumptions explicit before operational data are interpreted as evidence of comparative performance. The broader implication is that data useful for monitoring deployed AI systems are not automatically valid benchmarks for evaluating them.