发表机构
École polytechnique; Inria Paris; UC Berkeley(巴黎综合理工学院; 法国国家信息与自动化研究所巴黎分部; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生成器-验证器循环中错误接受累积问题,提出基于e值和共形风险控制的停止规则,控制错误发现率,并在合成数据和蛋白质设计基准上验证有效性。
AI 中文摘要
许多智能体工作流基于生成器-验证器循环:生成器提出候选,廉价验证器对其评分,当提案被验证为足够好时工作流终止。验证器通常代理更昂贵的真实标签预言机,当生成器针对其自适应搜索时,错误接受可能累积。提案可能通过代理检查但在更昂贵的真实标签检查中失败。我们研究何时停止这些循环,同时控制接受提案的错误发现率。我们的构造引入了在无分布统计检验和共形风险控制中具有独立兴趣的工具,包括通过索引投注构建的$e$-值分析,以及针对非单调损失的新型共形风险控制程序。我们在合成设置和蛋白质设计基准上验证了该方法。
英文摘要
Numerous agentic workflows are based on a generator-verifier loop: a generator proposes candidates, a cheap verifier scores them, and the workflow terminates when a proposal is verified as good enough. The verifier typically proxies a more costly ground-truth oracle, and as the generator searches adaptively against it, false acceptances may accumulate. Proposals can pass the proxy but fail under the costlier ground-truth check. We study when to stop these loops while controlling the false discovery rate of the accepted proposals. Our construction introduces tools of independent interest in distribution-free statistical testing and conformal risk control, including analysis of $e$-values constructed through index betting and a novel conformal risk control procedure for non-monotone losses. We validate the approach in synthetic settings and on a protein-design benchmark.