可用护栏:跨机器学习系统的选择性预测认证
Available Guardrails: Certifying Selective Prediction across ML Systems
浏览论文内容
中文总结 AI 辅助
本文提出通过精确二项反演和动态规划,使选择性预测的认证可用性可计算,并揭示有限数据恢复是核心挑战,从而将认证可用性作为可规划的部署资源。
中文摘要 AI 辅助
选择性预测器充当安全门:仅当预测看起来足够可信时才返回输出。部署中日益要求在目标精度下,对每个感兴趣的报告单元(如工具、策略标签或患者亚组)认证这种可靠性。主要困难通常不在于授予的证书是否有效,而在于有限的校准数据能否产生证书。随着安全门变得更安全或更细粒度,某些单元可能获得过少证据而无法认证。我们通过经典精确二项反演使这种可用性概念可计算,并在固定组顺序下将报告分区选择表述为动态规划,以揭示安全性、粒度和服务流量之间的权衡。由此产生的前沿揭示了一个有限样本估计几乎抹去的大规模人口机会:知情规划者在支持平衡上平均覆盖率提升0.157,而朴素估计器仅恢复0.005,使得从有限数据中恢复成为核心挑战。在一个规划分割上构建候选分区并在另一个分割上选择它们,可部分恢复这一差距,将平均覆盖率相对于支持平衡提升0.060,该方向在三个意图路由数据集和两个架构的60个模型效应中的59个中得到重现。一个互补的保持有效性的杠杆——跨报告单元重新分配族系误差预算——在人口数量和噪声估计下均能恢复额外覆盖率。相同的前沿在LLM工具调用、内容审核、病变分类和推荐中重复出现,且具有预测器特定的上限。因此,认证可用性是一种可规划的部署资源,它决定了安全门何时能被认证、以何种粒度认证以及覆盖多少流量。
英文摘要
A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficulty is often not whether a granted certificate is valid, but whether finite calibration data can produce one at all. As the gate becomes safer or more fine-grained, some units may receive too little evidence to certify. We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic. The resulting frontier reveals a large population opportunity that finite-sample estimation nearly erases: a truth-informed planner gains $0.157$ mean coverage over support balancing, whereas a naive estimator recovers only $0.005$, making recovery from finite data the central challenge. Constructing candidate partitions on one planning split and selecting among them on another recovers part of this gap, improving mean coverage over support balancing by $0.060$, with the direction reproduced in $59$ of $60$ model effects across three intent-routing datasets and two architectures. A complementary validity-preserving lever, reallocating the familywise error budget across reporting units, recovers additional coverage both with population quantities and noisy estimates. The same frontier recurs, with predictor-specific ceilings, across LLM tool-calling, content moderation, lesion classification, and recommendation. Certified availability is therefore a plannable deployment resource that determines when a safety gate can be certified, at what granularity, and over how much traffic.
发表机构
- Rivian and Volkswagen Group Technologies(Rivian与大众集团技术公司)
- Stony Brook University(石溪大学)
- Westlake University(西湖大学)
机构由 AI 辅助整理,请以论文原文为准。