arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11705cs.SE

哪种优化器,在何种预算下?基于搜索的软件演化的优化器竞赛

Which Optimizer, At What Budget? A Tournament of Optimizers for Search-Based SE

Kishan Kumar Ganguly, Tim Menzies

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对软件配置优化器选择问题,基于数据假设聚类20种优化器,通过竞赛对比其在不同标注预算下表现,发现最佳优化器随预算变化,还找到可替代大量计算的查表法,其预测效果良好且相关资源已开源。

中文摘要 AI 辅助

配置和调整现代软件是不可避免、成本高昂且容易出错的:一个系统可能有数百个相互作用的选项,评估一种设置可能意味着进行一次完整的构建或测试运行。标准的应对方法是自动化优化,但可用的优化器数量众多且不断增加。并且一些选择优化器的指导具有误导性:例如,NSGA-II被广泛推荐,但其他算法仅用其1/20的评估次数就能达到相同的结果。为帮助从业者更好地选择配置系统的工具,我们基于关于数据的六个假设对20种优化器进行聚类。接下来,我们在这些优化器之间进行了一场竞赛,使用106个软件演化优化任务,设置了四种标注预算(耗时超过14000个CPU小时)。我们发现没有一个优化器能完全胜出。最佳优化器会随着预算变化(从标签稀缺时的几何主动学习器变为标签丰富时的差分进化算法),所以在一种预算下“胜出”的优化器在另一些任务的另一种预算下可能是错误的。由于CPU成本,为每个新领域运行这样的竞赛是不切实际的。幸运的是,我们发现这14000小时的计算可以通过基于两个易于获取的任务属性(加上标注预算)的查表来替代。该表的预测在约75%的保留任务中与事后诸葛亮式的预测相当或更优。为支持开放科学,我们的竞赛和复制包已在这个https网址向软件演化搜索的研究人员和从业者开源。

英文摘要

Configuring and tuning modern software is unavoidable, expensive, and error-prone: a single system can expose hundreds of interacting options, and scoring one setting can mean a full build or test run. The standard response is automated optimization, but the number of available optimizers is large and growing. And some of the guidance for selecting among them is misleading: NSGA-II, for example, is widely recommended, yet other algorithms reach the same results using only 1/20th as many evaluations. To help practitioners make better choices about tools to configure their systems, we cluster 20 optimizers, based on six assumptions about the data. Next, we run a tournament across those optimizers, using 106 SE optimization tasks at four labeling budgets (taking 14,000+ CPU hours). We find that no optimizer wins outright. The best one migrates with the budget (from a geometric active learner when labels are scarce to differential evolution when labels are plentiful) so a winner "crowned" at one budget is wrong at another on up to half our tasks. Running such a tournament for every new domain is impractical due to its CPU cost. Fortunately, we find that those 14,000 hours can be replaced by a table lookup over two cheap-to-obtain task attributes (plus the labeling budget). Predictions from this table tie or beat a hindsight oracle on $\approx 75%$ of held-out tasks. To support open science, our tournament and replication package are open-sourced for SBSE researchers and practitioners at https://github.com/KKGanguly/OptimizerTournament.

↑