基于不完美代理奖励的可靠自进化
Reliable Self-Evolution with Imperfect Proxy Rewards
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对自进化搜索中代理奖励不完美导致假阳性污染的问题,提出共形区间驱动的自进化(CISE)方法,通过统计校准的奖励区间和保守反馈,在材料科学任务中实现高保真评估下全部真阳性,验证了更小但更精确候选清单的价值。
AI中文摘要:
基于大型语言模型(LLM)的自进化搜索是科学发现的一种有前景的方法。然而,在某些领域,对每个候选方案进行高保真评估的成本高得令人望而却步。因此,此类环境中的自进化系统依赖于低成本但不完美的代理奖励,这些奖励可能会给不可行的候选方案打高分。这些假阳性可能会污染最终输出以及用于指导后续生成的反馈。这促使我们采用统计校准的奖励区间来实现更可靠的自进化搜索。我们提出了共形区间驱动的自进化(CISE),该方法利用条件共形推断和逐轮在线密度比估计来构建候选特定的奖励区间。CISE 在进化反馈中使用保守的基于区间的奖励,并且仅当所有所需属性区间完全位于其各自的可行区域内时才返回候选方案。我们在独立性和协变量偏移的明确假设下推导了固定迭代的覆盖结果。我们在材料科学的三个自进化搜索任务上评估了 CISE。在我们的实验中,CISE 返回的所有候选方案在高保真评估下均为真阳性,而基线方法返回的候选方案更多但包含假阳性。这些结果凸显了在下游验证预算有限时,更小但更精确的候选清单的价值。我们的代码库可从此 https URL 获取。
英文摘要:
Large language model (LLM)-based self-evolving search is a promising approach to scientific discovery. However, high-fidelity evaluation of every candidate is prohibitively expensive in some domains. Self-evolving systems in such settings therefore rely on low-cost but imperfect proxy rewards, which may assign high scores to infeasible candidates. These false positives may contaminate both the final output and the feedback used to guide subsequent generations. This motivates statistically calibrated reward intervals for more reliable self-evolving search. We propose Conformal Interval-Driven Self-Evolution (CISE), which constructs candidate-specific reward intervals using conditional conformal inference and iteration-wise online density-ratio estimation. CISE uses conservative interval-based rewards for evolutionary feedback and returns candidates only when all required property intervals lie entirely within their respective feasible regions. We derive fixed-iteration coverage results under explicit assumptions of independence and covariate shift. We evaluate CISE on three self-evolving search tasks in materials science. In our experiments, all candidates returned by CISE are true positives under high-fidelity evaluation, whereas the baselines return more candidates but include false positives. These results highlight the value of a smaller, more precise shortlist when downstream validation budgets are limited. Our repository is available at https://github.com/MLAI-Yonsei/CISE.git.