arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

认证或拒绝:协变量偏移下带覆盖下限的选择性风险控制的跨模型映射

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

Jiamiao Liu, Dewen Qiao, Yu Zhang, Xuetao Chen

arXiv 2608.10893首次发表:更新:

AI 中文总结

该研究针对协变量偏移下的选择性风险控制,提出带覆盖下限的跨模型映射,给出Model-A、Model-B、Model-B'的相关界,并通过实证验证了其有效性。

AI 中文摘要

认证选择性预测器能达到其所能达到的任何覆盖率;而操作人员设定了自动化下限:在协变量偏移的情况下,至少要回答β比例的偏移目标流量,且错误回答的比例不超过α。在有界比例协变量偏移下,我们证明了存在“下限认证映射”:一旦必须在认证选择条件风险α的同时对该下限进行认证,认证就会获得一个可行性边界和一个双资源复杂度映射,其复杂度可加至常数项,涉及带标签源数据的风险、无标签目标样本的下限。该速率是局部的,需要规则的边界裕度、局部区域阈值以下的松弛度以及格点条件:上界需预先注册格点裕度,下界需与各松弛度兼容。所展示的划分是操作路径;神谕权重还允许对带标签源数据的下限进行估计。三项带模型标签的结果:下界(Model-B)、匹配神谕权重的上界(Model-A),以及在预先注册的精确分层偏移模型下有效的可实现上界(Model-B'),该模型明确计入了干扰成本。这种匹配是跨这些模型而非单模型极小极大定理实现的,且必须如此:在整个有界比例类别中,任何未知权重的过程在任何样本量下都无法匹配(Model-B不一致,在α=β=1/2时可验证)。干扰的必要性仅部分得到解决。复杂度在两侧均跟踪局部化接受区域的泛函,而非全局有效样本量(ESS),尽管固定ESS分离定理仍未解决;当β趋近于0时,两个下界轴均消失,因此下限构成了该映射。实证方面,注册的“ bite”族在其预先注册的区间内呈现出-2.002的双对数斜率;1024个单元的审计记录到0次违规,而正式认证已生效;单语料库SQuAD到NewsQA的可行性审计返回了诚实的弃权(不执行)。

英文摘要

Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a $β$-fraction of shifted target traffic with at most an $α$-fraction of answers wrong. Under bounded-ratio covariate shift we prove the Floor Certification Map: once that floor must be certified alongside the selection-conditioned risk $α$, certification acquires a feasibility frontier and a two-resource complexity map, additive up to constants: risk in labeled source, the floor in unlabeled target samples. The rates are local, needing a regular frontier margin, slack below the local-regime threshold, and lattice conditions: pre-registered with a lattice margin for the upper bounds, compatible per-slack for the lower. The displayed split is the operational route; oracle weights also allow a labeled-source floor estimate. Three model-tagged results: a lower bound (Model-B), a matching oracle-weight upper bound (Model-A), and an implementable upper bound (Model-B') valid under a pre-registered exact stratified-shift model with nuisance cost priced explicitly. The match is across these models rather than a single-model minimax theorem, and necessarily so: over the full bounded-ratio class no unknown-weight procedure matches at any sample size (Model-B is inconsistent, witnessed at $α=β=1/2$). The nuisance's necessity is only partially settled. Complexity tracks a localized accepted-region functional, not global effective sample size (ESS), on both sides, though a fixed-ESS separation theorem is left open; both lower-bound axes vanish as $β\to0$, so the floor creates the map. Empirically, the registered bite family diverges with log-log slope $-2.002$ within its pre-registered band; a 1,024-cell audit records 0 violations where the formal certificates fire; and a single-corpus SQuAD-to-NewsQA feasibility audit returns honest refusal.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑