arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReMCTS:用于代码生成的反射增强蒙特卡洛树搜索

ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation

Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan

arXiv 2609.34717首次发表:更新:

发表机构

Shenzhen University(深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReMCTS通过执行引导、记忆增强的MCTS式搜索,在代码生成中利用跨分支失败经验,提升HumanEval和MBPP上的生成质量。

AI 中文摘要

开放权重的大型语言模型(LLMs)能够根据自然语言提示生成函数级程序,但看似合理的候选程序仍会在隐藏语义上失败,并在多次修复尝试中重复错误。我们提出了ReMCTS,一种基于执行、记忆增强、LLM引导的MCTS式搜索框架。它将程序候选组织为树状态,保留分支局部的调试上下文,跨分支检索失败经验,并区分失败的检查与不可用的证据。在HumanEval和MBPP-Sanitized上,在保留评估下,可见测试的ReMCTS在10个模型-数据集配对中的8个上优于直接生成,而仅使用代理的搜索则较不稳定。受控的树搜索、采样、修复和记忆消融实验刻画了这些收益的来源和局限。一个30任务的HumanEval-X C++试点进一步证明了与编译器支持的执行的兼容性,但并未构成广泛的多语言评估。

英文摘要

Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.

Comments21 pages, 2 figures. To appear in the Proceedings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑