Ockhamareto:用于强化学习生成简洁单元测试的帕累托门控片段级信用分配
Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
Ockhamareto是基于奥卡姆剃刀和帕累托最优的GRPO框架,在多个基准上优于MIST-RL等方法,可生成更简洁高效的单元测试,适用于不同规模模型。
中文摘要 AI 辅助
我们提出Ockhamareto,这是一种基于奥卡姆剃刀和帕累托最优原则的单次GRPO框架,用于单元测试的生成与选择。Ockhamareto包含两个核心组件:(i)帕累托门控奖励,仅奖励在(变异,-测试数量)空间中非支配的rollout;(ii)令牌级片段信用,将每个测试的边际变异杀死归因于其单元测试块的令牌。在UnLeakedTestBench(ULT)上,Ockhamareto严格帕累托支配最强的RL基准MIST-RL,且在所有优化目标上均占优:在N=5时,变异得分从31.3%提升至49.9%,使用的测试数量从平均4.67个减少至2.60个,实现了3.4倍的每测试权衡改进。在HumanEval+、MBPP+、CodeContests、TestGenEval-Lite这四个基准上,Ockhamareto在所有基准中均领先于变异和覆盖度指标,且始终拥有最小的测试套件;在4B、9B、27B等所有模型规模下,其变异得分均比当前最优方法高出30-35个百分点。我们还发现,帕累托前沿上效率与有效性最优权衡的膝点,与函数大小等易计算的代理指标无明显相关性,这一发现推动了帕累托前沿计算的必要性,需为每个被测函数确定这一关键工程权衡。
英文摘要
We introduce \textbf{Ockhamareto}, a single-shot GRPO framework for unit-test generation and selection, based on the principles of \emph{Ockham's Razor} and \emph{Pareto Optimality}. Ockhamareto has two principal components: (i)~a \emph{Pareto-gated Bonus} that rewards only rollouts non-dominated in~(mutation, $-$\#tests) space, and (ii)~\emph{Token-level Segment Credit}, which attributes each test's marginal mutation kills back to the tokens of its unit-test block. On the \emph{UnLeakedTestBench~(ULT)}, Ockhamareto \emph{strictly Pareto-dominates} the strongest RL baseline~(\emph{MIST-RL}). Furthermore, it dominates on {\em each and all} optimization objectives, catching more bugs ($49.9\%$ vs $31.3\%$ mutation score at $N{=}5$), using \emph{fewer} tests ($2.60$ vs $4.67$ on average), thereby achieving $3.4\times$ the per-test trade-off improvement. The advantage is found in all four benchmarks~(\emph{HumanEval+}, \emph{MBPP+}, \emph{CodeContests}, \emph{TestGenEval-Lite}): Ockhamareto leads both mutation and coverage metrics on every one, always with the smallest suite. Ockhamareto also outperforms the state-of-the-art at all model scales, adding $+30$--$35$~pp mutation at 4B, 9B, and 27B model sizes. We also show that the knee point of the optimal trade-off between efficiency and effectiveness on the Pareto front is not correlated with obvious more easily computed proxy metrics, such as function size. This finding motivates the Pareto front computation; it is needed to identify this crucial engineering trade-off for each function under test.
发表机构
- National University of Singapore(新加坡国立大学)
- University College London(伦敦大学学院)
- King’s College London(伦敦国王学院)
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。