AI 中文总结
本文提出评估盲区概念,指出测量故障会从AI训练到部署阶段无声传播,通过案例与50起真实事件验证,发现53%公开失效为无声,强调测量基础设施需覆盖全生命周期。
AI 中文摘要
AI系统可能会无声地失效,这种失效会在训练循环、评估流程和生产监控栈中传播,直到下游产生危害才变得可见。本文引入了评估盲区(evaluation blindness)的概念:当测量函数M在系统实际失效时,产生的读数与健康状态无法区分,且没有辅助信号标记这种差距时,M相对于失效类别F就表现出评估盲区。该问题出现在文献分开处理的两个生命周期阶段:在训练阶段,奖励模型被操纵、重要性采样校正被无声地错误计算、基准污染会夸大微调评估,而损失曲线看起来正常,梯度更新也正常进行;在部署阶段,监控未能捕获六类生产失效,包括结构定义上100%无声的操作(Operational)类别。我们提供了一个统一两个阶段的形式可检测性谓词。四项训练阶段案例研究追踪了具体的故障,包括TRL PR #6594中的一个真实实现bug,其中梯度在损失正常下降时被破坏。针对来自法院文件和监管备案的50起真实事件验证的六类分类法发现,53%可验证的公开失效是无声的。失效预算框架将可接受的失效率与用例风险类别关联起来。其含义直接:测量基础设施是整个AI生命周期中的正确性问题,而不仅仅是在评估时。数据、代码和分类法架构在此https URL。
英文摘要
AI systems can fail silently. The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream harm makes it visible. This paper introduces evaluation blindness: a measurement function M exhibits evaluation blindness with respect to failure class F when it produces readings indistinguishable from a healthy state while the system is actually failing, with no auxiliary signal flagging the gap. The problem surfaces at two lifecycle stages the literature has treated separately. At training time, reward models are gamed, importance-sampling corrections are silently miscalculated, and benchmark contamination inflates fine-tuning evaluations, all while loss curves look healthy and gradient updates proceed normally. At deployment time, monitoring fails to catch six classes of production failure, including an Operational category that is 100% silent by structural definition. We provide a formal detectability predicate unifying both stages. Four training-time case studies trace concrete breakdowns, including a real implementation bug in TRL PR #6594 where gradients are corrupted as loss decreases normally. A six-class taxonomy validated against 50 real-world incidents from court documents and regulatory filings finds that 53% of verifiable public failures were silent. A failure budget framework ties acceptable failure rates to use-case risk class. The implication is direct: measurement infrastructure is a correctness concern across the full AI lifecycle, not just at evaluation time. Data, code, and taxonomy schema are at https://github.com/priyanka25aug/llm-failure-taxonomy.
Comments10 pages, 7 figures