arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21748cs.CV

校准你所发布的内容:验证器引导的文本到图像生成的后选择风险控制

Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation

Xuanhua Yin, Shunqi Mao, Wei Guo, Chuanzhi Xu, Weidong Cai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对验证器引导的文本到图像生成的后选择风险控制问题,提出 SHIP 方法,通过策略级校准降低发布输出风险,在 GenEval2 数据集上使发布风险从 0.310 降至 0.162。

中文摘要 AI 辅助

验证器引导的文本到图像(T2I)系统越来越多地在测试时使用搜索来从多个候选中选择、细化或停止生成,但发布的阈值通常是针对单个图像校准的。这会造成候选到策略的校准不匹配:搜索既改变了接收输出的提示,也改变了要发布的候选,因此候选级别的风险控制不一定意味着对发布输出的风险进行控制。我们通过提示重加权和提示内选择来形式化这种估计量偏移,并提出 SHIP(Selection-aware Held-out calibration of Inference Policies,即感知选择的推理策略保留校准)。SHIP 在保留的提示上运行或重放完整的部署策略,使用独立的目标评判器评估其实际发布的图像,并选择最宽松的阈值,其风险上界满足规定的预算。对于具有预先指定阈值网格的可重放策略,同时置信控制提供有限样本有效性。在固定、顺序和自适应 T2I 推理过程上的实验表明,策略级校准可恢复更低风险的操作点,同时揭示风险、覆盖率和计算之间的策略依赖权衡。在 GenEval2 数据集上使用 FLUX 模型,当 N=16 时,合并候选阈值产生的发布风险为 0.310,而 SHIP 将其降低至 0.162。在 200 个缓存流划分上,固定网格证书无目标越界。因此,可靠的推理时扩展需要校准完整部署策略所诱导的输出分布。

英文摘要

Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds are often calibrated on individual images. This creates a candidate-to-policy calibration mismatch: search changes both which prompts receive an output and which candidate is released, so candidate-level risk control need not imply control of released-output risk. We formalize this estimand shift through prompt reweighting and within-prompt selection, and introduce SHIP, Selection-aware Held-out calibration of Inference Policies. SHIP runs or replays the complete deployed policy on held-out prompts, evaluates the image it actually releases using an independent target judge, and selects the most permissive threshold whose risk upper bound satisfies a prescribed budget. For replayable policies with a prespecified threshold grid, simultaneous confidence control provides finite-sample validity. Experiments across fixed, sequential, and adaptive T2I inference procedures show that policy-level calibration recovers lower-risk operating points while exposing policy-dependent tradeoffs among risk, coverage, and compute. On GenEval2 with FLUX at N=16, a pooled-candidate threshold yields released risk 0.310, whereas SHIP reduces it to 0.162. Across 200 cached-stream splits, the fixed-grid certificate has no target crossing. Reliable inference-time scaling therefore requires calibrating the output distribution induced by the complete deployed policy.

补充信息

↑