AI 中文总结
Janus是一个算法-评估器协同进化框架,它用LLM协同进化目标程序与代理评估器,通过真实结果校准、在线信用更新等策略缓解分布偏移,在五项任务中用更少真实评估实现更优性能,将评估器引导的LLM发现扩展到评估昂贵的科学领域。
AI 中文摘要
大语言模型(LLM)驱动的程序发现依赖于快速的评估器反馈,但许多科学和工程任务需要高保真模拟、硬件执行或物理实验,这使得每次评估的成本很高。廉价的替代评估器可以降低成本,但固定的替代评估器容易受到搜索诱导的分布偏移影响,且难以从稀疏、搜索偏向的标签中可靠拟合。我们引入Janus,一个使用LLM协同进化目标程序和可执行代理评估器的框架。为解决标签稀缺问题,Janus利用LLM中编码的领域知识生成任务特定的评估器程序,并使用真实结果对其进行校准。为缓解分布偏移,Janus与目标程序协同进化评估器,采用与晋升对齐的目标来选择评估器,并维护带有在线信用更新的区域条件组合。由于代理预测仍可能出错,Janus仅用它们来优先选择候选者,且要求候选者在进入目标程序群体或更新当前最优解前必须经过真实验证。在五项科学和工程设计任务中,Janus在真实评估预算下的最佳至今改进曲线下面积,比仅进化目标程序的匹配基线更大,最终性能也更高。平均而言,Janus使用少59.1%的真实评估,达到基线最终改进的99%。进化后的代理评估器也比其种子版本更准确地对有前景的候选者进行排名。这些结果共同将评估器引导的LLM发现,从具有廉价、可扩展反馈的任务,扩展到可信评估稀缺且昂贵的科学领域。
英文摘要
LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive. Cheap surrogate evaluators can reduce this cost, yet fixed surrogates are vulnerable to search-induced distribution shift and are difficult to fit reliably from sparse, search-biased labels. We introduce Janus, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators. To address label scarcity, Janus leverages domain knowledge encoded in LLMs to generate task-specific evaluator programs and calibrates them using real outcomes. To mitigate distribution shift, Janus evolves evaluators alongside target programs, selects them using a promotion-aligned objective, and maintains region-conditioned portfolios with online credit updates. Because proxy predictions remain fallible, Janus uses them only to prioritize candidates and requires real validation before candidates can enter the target-program population or update the incumbent. Across five scientific and engineering design tasks, Janus achieves a larger area under the best-so-far improvement curve over the real-evaluation budget and higher final performance than a matched baseline that evolves only target programs. On average, Janus reaches 99/% of the baseline's final improvement with 59.1/% fewer real evaluations. Evolved proxy evaluators also rank promising candidates more accurately than their seed versions. Together, these results extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.