发表机构
King’s College London, University of London; National University of Singapore; University College London, University of London(伦敦大学国王学院(伦敦大学); 新加坡国立大学; 伦敦大学学院(伦敦大学))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出感知难度的 COMPAS 方法,通过联合搜索模型、提示词和解码设置优化代码生成,在 LiveCodeBench 和 SWE-bench 上均实现性能提升且成本降低。
AI 中文摘要
代码生成系统每次调用大语言模型(LLM)时,都会使用一个模型、一个提示词(prompt)和解码设置。然而,现有的优化方法通常仅调整这些选择中的一部分,或对所有任务使用单一固定配置:全局优化器为所有任务搜索一个配置,路由模型仅选择模型,提示词优化器则保持模型和解码设置固定。这使得它们的联合配置、针对特定任务组的交互机制尚不明确。因此,本文研究这些选择之间的交互作用,观察到提示词和解码设置存在交互,调优效果因模型而异,且最佳配置随任务难度变化。基于这些观察,本文提出 COMPAS(全称:Code-generation Optimization over Models, Prompts, And Decoding Settings,即针对模型、提示词和解码设置的代码生成优化),这是一种感知难度的方法,通过低成本的模型选择和联合提示词-解码搜索,学习特定任务组的质量-成本前沿,随后在线将每个测试任务路由到其匹配的前沿,无需进一步搜索。在 LiveCodeBench 上使用匹配的搜索预算时,COMPAS 将 pass@1 指标从最佳基准的 45.9% 提升至 52.8%,同时将成本从 36.57 美元降至 4.92 美元。该方法还可迁移到 SWE-bench 上的仓库级代码生成任务,解决 76.0% 的任务,而最佳基准的解决率为 70.0%。代码和可复现的人工制品可在该 https URL 获取。
英文摘要
Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one fixed configuration for all tasks: global optimizers search one configuration for all tasks, routers choose only a model, and prompt optimizers keep the model and decoding settings fixed. This leaves their joint, group-specific interactions unclear. We therefore examine how these choices interact and observe that prompts and decoding settings interact, tuning effects vary by model, and the best configuration varies by task difficulty. Guided by these observations, we introduce COMPAS (Code-generation Optimization over Models, Prompts, And Decoding Settings), a difficulty-aware method that learns group-specific quality-cost fronts through low-cost model selection and joint prompt-decoding search, then routes each test task to its matching front online without further search. Under a matched search budget on LiveCodeBench, COMPAS improves pass@1 from 45.9% for the best baseline to 52.8% while reducing cost from $36.57 to $4.92. This also transfers to repository-level code generation on SWE-bench, resolving 76.0% of tasks versus 70.0% for the best baseline. Code and the reproducibility artifact are available at https://github.com/gjz78910/COMPAS.