发表机构
Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Narcissus是一种程序合成器,通过保留LLM提议的语法树并结合上下文评分与正则化项,在多领域搜索中优于静态引导及修复提议的方法,解决ARC任务的比例显著提升。
AI 中文摘要
大语言模型(LLMs)擅长编程,但当任务固定目标语言时则表现不佳:在训练数据中遇到罕见语法时,其生成的程序通常会违反语法或无法满足给定规范。枚举式合成器会系统地搜索语法正确的程序空间,并由LLMs引导;现有技术通过将LLMs的提议近似为规则频率来引导,这会丢失每个构造的所属位置,且会在提议出错时修剪掉所有提议未覆盖的规则。我们提出Narcissus,一种合成器,它将提议保留为语法树,并在候选程序的上下文中对每个扩展进行评分:具有相同周围结构的提议是否以相同方式延续,以及扩展是否重建了提议重复的片段。正则化项确保每个规则都可达,因此错误的提议只会延迟解决方案,而不会隐藏它。在五个领域和两个搜索后端上,Narcissus在每个预算下都优于静态引导,且始终胜过重新提示LLM以修复其自身提议的方法;它能快一个数量级地达到类似提议的程序,并解决了40%的ARC任务,而原始提议仅能解决13%,且在搜索过程中无需一次LLM调用。
英文摘要
Large language models (LLMs) excel at programming, but not when the task fixes the target language: prompted with a grammar rare in their training data, their programs usually break the grammar or fail the given specification. Enumerative synthesizers search the space of syntactically correct programs systematically guided by LLMs; the state of the art guides them by approximating LLM proposals into rule frequencies, which loses where each construct belongs and prunes every rule the proposals miss, exactly when the proposals are wrong. We present Narcissus, a synthesizer that keeps the proposals as syntax trees and scores each expansion of a candidate program in its context: does a proposal with the same surrounding structure continue the same way, and does the expansion rebuild a fragment the proposals repeat? A regularization term keeps every rule reachable, so wrong proposals delay the solution but cannot hide it. Across five domains and two search backends, Narcissus beats static guidance at every budget and consistently outperforms re-prompting the LLM to fix its own proposals; it reaches proposal-like programs an order of magnitude sooner and solves $40\%$ of ARC tasks where the raw proposals solve $13\%$, all without a single LLM call during search.