发表机构
Peking University; Alibaba Group; Zhejiang University(北京大学; 阿里巴巴集团; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MIMIC框架,利用可执行代码合成推理数据,通过代码插桩奖励提供过程监督,显著提升LLM在通用推理、数学和细粒度任务上的准确率。
AI 中文摘要
大型语言模型(LLMs)在编程任务上表现出色,但在自然语言中的确定性、细粒度推理上却经常失败,它们严重依赖语义近似而非稳健的符号执行。为弥合这一差距,我们提出了MIMIC框架,该框架利用可执行代码作为推理数据合成的严谨媒介。MIMIC通过叙事融合、代码引导的测试合成和动态代码插桩,将算法从根本上转化为可验证的推理轨迹。至关重要的是,这些显式的中间执行状态自然形成了一种代码插桩奖励(CIR),为强化学习提供了密集、高保真的过程监督,而无需外部奖励模型。大量评估表明,在我们合成的数据集上通过SFT和GRPO训练的模型取得了显著且一致的提升。我们的方法显著提高了通用推理、复杂数学基准以及细粒度确定性任务的准确率,证明可执行代码的程序严谨性能够有效释放并增强LLMs的泛化推理能力。我们的代码和数据可在以下网址获取:此https URL。
英文摘要
Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge this gap, we propose MIMIC, a framework that leverages executable code as a rigorous medium for reasoning data synthesis. MIMIC fundamentally transforms algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation. Crucially, these explicit intermediate execution states naturally form a Code-Instrumented Reward (CIR), providing dense, high-fidelity process supervision for reinforcement learning without external reward models. Extensive evaluations reveal that models trained via SFT and GRPO on our synthesized dataset achieve substantial, consistent gains. Our method significantly elevates accuracy across general reasoning, complex mathematical benchmarks, and fine-grained deterministic tasks, demonstrating that the procedural rigor of executable code can effectively unlock and enhance the generalized reasoning capabilities of LLMs. Our code and data are available at https://github.com/zjy1298/MIMIC.
CommentsAccepted by EMNLP26 main