arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于代码偏好优化的函数级执行反馈

Function-Level Execution Feedback for Code Preference Optimization

Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong, Jungwoo Lee

arXiv 2608.23632首次发表:更新:

发表机构

Seoul National University; Konkuk University(首尔大学; 建国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出STEP-KTODER框架,将代码步骤定义为多函数程序的模块级函数,结合函数级过程监督与程序结果级反馈,在多个代码基准测试中优于KTO、DPO等方法,且验证了执行标签的重要性。

AI 中文摘要

过程监督已提升了数学推理能力,其中间步骤自然表现为思维链。然而在代码生成领域,过程监督仍未得到充分探索,因为缺乏标准的步骤定义。监督可针对代码行、推理轨迹或程序状态,导致难以确定标记和优化的对象。我们提出STEP-KTODER,这是一个代码偏好优化框架,将步骤定义为分解后的多函数程序中的模块级函数,并通过自动生成的单元测试分配二元正确性标签。该方法是分步KTO的代码特定实例,结合了函数级过程监督与完整程序的结果级反馈。我们在HumanEval(+)、MBPP(+)、BigCodeBench和LiveCodeBench上进行评估,结果显示STEP-KTODER优于仅基于结果的KTO和DPO。进一步分析表明,基于执行的标签至关重要:大语言模型(LLM)作为评判者的注释会系统性地高估函数失败情况,破坏正向步骤标签,并降低下游偏好优化的效果。代码可从以下网址获取:this https URL。

英文摘要

Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In code generation, however, process supervision remains underexplored because there is no standard notion of a step. Supervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize. We propose STEP-KTODER, a framework for code preference optimization that defines steps as module-level functions in decomposed multi-function programs and assigns binary correctness labels via automatically generated unit tests. Our method provides a code-specific instantiation of stepwise KTO, combining function-level process supervision with outcome-level feedback on the full program. We evaluate on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench, showing that STEP-KTODER improves over outcome-only KTO and DPO. Further analysis shows that execution-based labels are essential: LLM-as-a-judge annotations systematically over-predict function failures, corrupt positive step labels, and degrade downstream preference optimization. Code is available at: https://github.com/inechnech/STEP-KTODER.

Comments20 pages, 8 figures, 14 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑