AI验证的几何:独立同分布N选最优搜索的精确认证极限
The geometry of AI validation: From structural blindness to reusable audits
- Technical University of Darmstadt(达姆施塔特工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对独立同分布N选最优搜索推导了精确认证极限,提出双门审计规则,其在数学推理和代码选择的回顾性分析中可降低保留误差。
AI中文摘要:
AI系统越来越多地生成备选方案、检查证据并部署选定的输出,因此验证具有目标相关性:只有在产生该证据的干预所解决的方向上,证据才能证明部署的合理性。我们将验证和部署规则表示为可靠性曲面上的核函数,其跨度几何将减少采样噪声的复制与减少结构盲区的新干预方向区分开来。我们针对独立同分布(iid)N选最优搜索精确实现了这一原理。在标量排序、随机平局、最大选择、有界二元真值及稳定的真值-秩关系下,通过n=m个叶节点确定N选最优的可靠性时,精确歧义宽度为B_{m,N}=1+2∑_{r=1}^{m}(-1)^r cos^{2N}[rπ/(2(m+1))]。显式有界区间可达到整个区间,且在限定n≤m的可靠性均值审计中,完整前缀具有信息最大化特性。控制尺度为m²/N:当m与√N成比例时,歧义保持约0.83;而宽度ε要求m的量级为√[N log(1/ε)]。单调性给出精确的均匀近似边界;利普希茨界给出精确的截尾对偶和阶精确的L/m²歧义。这些结果产生了双门审计规则:先建立结构覆盖,再添加独立任务以提高精度。对数学推理和代码选择的回顾性研究构建了具有宽间隔的兼容部署值,显示冻结在82个发现任务上的分数尾审计规则大幅降低了保留误差。超出iid搜索范围后,该几何仅适用于已知或独立估计的核函数;实证分析为说明性而非前瞻性干预。
英文摘要:
AI systems increasingly search among candidate answers and deploy the highest-scoring one. Increasing search changes which errors matter, so a precise evaluation at one computation budget can leave another budget unresolved. We connect this information gap to the cost of closing it. For independent best-of-n search, aggregate reliability measurements identify deployment only through the directions they observe; we derive an exact ambiguity frontier when only small search widths are audited. Retaining candidate ranks and truth labels enables a constructive alternative: one audit can estimate reliability across all widths up to N. With known score percentiles, the minimax worst-coordinate mean squared error scales as (1 + log N)/T + N/M, capped at a constant, for expected budgets of T truth labels and M candidate observations. Matching lower bounds allow adaptive label acquisition, establishing that the distinct label and candidate costs are intrinsic to this experiment. An explicit design attains this order; a complementary record-based procedure supplies simultaneous guarantees without a known score distribution. Retrospective mathematical-reasoning and code-generation analyses show why search-dependent validation matters. In held-out CodeRM pools, a shared audit reduces the 95th-percentile maximum error across 100 widths by 58% and 40% relative to uniform labeling. These results turn structural ambiguity into a quantitative prescription for reusable validation.