发表机构
Nankai University; University of Waterloo(南开大学; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM自动启发式设计中求解器修订的可行性振荡问题,提出DiSH控制机制,联合控制修订方向与范围,在九个CO基准上提升QYI分数并减半反馈消耗。
AI 中文摘要
Harness(控制机制)正通过使大语言模型(LLM)智能体的行为更可控来日益释放其潜力。在面向组合优化(CO)的自动启发式设计(AHD)中,LLM利用执行和验证反馈迭代地修订求解器。然而,这些修订常常破坏先前已验证的求解器行为,使得优化过程不稳定且成本高昂。我们将这种模式刻画为可行性振荡,并在多个CO基准、AHD方法和骨干LLM中记录了该现象。追踪这一振荡,我们的分析揭示,可行性退化往往紧随大规模的求解器修订,而大量的反馈冗余使得识别哪些失败值得关注变得更加困难。受这些发现启发,我们提出了DiSH,一种联合控制修订方向和范围的harness(控制机制)。它选择具有代表性的失败证据来指导每次更新,并随着可行行为的积累逐步缩小修订范围。在九个CO基准上,DiSH一致地提高了QYI分数并减少了可行性振荡,同时与使用完整反馈的先前方法相比,反馈令牌消耗大约减半。这些结果确立了DiSH作为harness控制下求解器优化的强基线。
英文摘要
Harnesses are increasingly unlocking the potential of large language model (LLM) agents by making their actions more controllable. In automatic heuristic design (AHD) for combinatorial optimization (CO), LLMs iteratively revise solvers using execution and verification feedback. However, these revisions often break previously validated solver behavior, making refinement unstable and costly. We characterize this pattern as feasibility oscillation and document it across CO benchmarks, AHD methods, and backbone LLMs. Tracing this oscillation, our analysis reveals that feasibility degradation often follows large solver revisions, while substantial feedback redundancy makes it harder to identify which failures deserve attention. Motivated by these findings, we introduce DiSH, a harness that jointly controls revision direction and scope. It selects representative failure evidence to guide each update and progressively narrows the revision scope as feasible behavior accumulates. Across nine CO benchmarks, DiSH consistently improves QYI scores and reduces feasibility oscillation, while roughly halving feedback token consumption relative to prior methods that use full feedback. These results establish DiSH as a strong baseline for solver refinement under harness control.
CommentsThe authors are listed in alphabetical order