发表机构
Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM代码生成的静态绑定问题,提出动态上下文适配方法,结合验证智能体、知识图谱与模拟退火,在8个问题中7个上优于零样本、Reflexion等方法,1000次评估时在交叉耦合优化任务也表现最佳。
AI 中文摘要
基于大语言模型(LLM)的代码生成在正确性依赖于执行相关耦合时会失效:一个例程的含义由另一个例程的运行时行为定义,这种关系无法仅通过文本描述解决,我们将此限制称为静态绑定。该限制不仅存在于显式耦合问题中,当正确性依赖于组件间的联合执行行为时,从显式交叉耦合优化器到打包、路由和符号搜索中更微妙的联合约束,它都会以不同程度出现。本文提出动态上下文适配,这是一种为该场景设计的样本高效验证-生成循环:验证智能体从执行轨迹中提取结构化诊断信息,为生成智能体提供类梯度的指导,生成智能体每次迭代提出多个候选;源自问题描述的知识图谱为生成智能体提供语义约束;模拟退火在候选中进行选择以避免贪婪坍缩。在8个问题中的7个上,无论是300次评估还是600次评估,我们的方法都优于零样本、Reflexion和OpenEvolve(p < 0.01),在该 regime 中,基于种群的搜索尚未积累足够多样性以竞争。值得注意的是,在作为主要动机问题的交叉耦合优化上,我们的方法在1000次评估时也取得了最佳分数,这与结构化执行反馈在正确性依赖于运行时耦合时最有益的假设一致。消融实验结果证实,结构化执行反馈是主要驱动因素。
英文摘要
LLM-based code generation fails when correctness depends on execution-dependent coupling: the meaning of one routine is defined by the runtime behavior of another, a relationship that cannot be resolved from textual descriptions alone. This limitation, which we call static binding, is not confined to explicitly coupled problems; it appears to varying degrees whenever correctness depends on joint execution behavior across components, from explicit cross-coupled optimizers to subtler joint constraints in packing, routing, and symbolic search. This paper proposes dynamic context adaptation, a sample-efficient validation-generation loop designed for this setting. A validation agent extracts structured diagnostic information from execution traces, providing gradient-like guidance to a generation agent that proposes multiple candidates per iteration. A knowledge graph derived from the problem description supplies semantic constraints to the generation agent. Simulated annealing selects among candidates to avoid greedy collapse. Our method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), a regime where population-based search has not yet accumulated sufficient diversity to compete. Notably, on the primary motivating problem (cross-coupled optimization), our method also achieves the best score at 1000 evaluations, consistent with the hypothesis that structured execution feedback is most beneficial when correctness depends on runtime coupling. Ablation results confirm that structured execution feedback is the primary driver.
CommentsAccepted at ICANN 2026