AI 中文总结
针对现有论文到代码复现系统因静态规划易出现不一致的问题,提出DeepRepro状态感知框架,通过动态生成子规划并结合编排与监控,在PaperBench Code-Dev上表现优于基线
AI 中文摘要
智能体大语言模型(LLMs)的近期进展已实现日益自主的软件工程工作流,但自动机器学习(ML)论文到代码的复现仍是极具挑战性的长时序问题。与传统代码生成不同,该任务需构建并维护一个功能完备的仓库,其状态在执行过程中持续演化。现有系统通常依赖静态的前期规划,随后进行逐文件生成,这常因依赖项、接口及执行反馈随时间变化而导致不一致。我们提出DeepRepro,一种基于执行状态感知子规划的论文到代码复现框架。DeepRepro将演化的仓库状态与运行时反馈动态转化为细粒度的实现子规划,使规划在整个仓库构建过程中与执行保持一致。该框架还融入了仓库感知编排和轻量级过程感知接口,用于透明监控长时序复现。在PaperBench Code-Dev上的实验表明,DeepRepro始终优于强大的科学与商业代码智能体基线。
英文摘要
Recent advances in agentic large language models (LLMs) have enabled increasingly autonomous software engineering workflows, yet automatic machine learning (ML) paper-to-code reproduction remains a challenging long-horizon problem. Unlike conventional code generation, this task requires constructing and maintaining a fully functional repository whose state continuously evolves during execution. Existing systems typically rely on static upfront planning followed by sequential file-level generation, which often leads to inconsistencies as dependencies, interfaces, and execution feedback change over time. We propose DeepRepro, a state-aware framework for paper-to-code reproduction based on execution-state-aware subplanning. DeepRepro dynamically transforms evolving repository states and runtime feedback into fine-grained implementation subplans, keeping planning aligned with execution throughout repository construction. The framework further incorporates repository-aware orchestration and a lightweight process-aware interface for transparent monitoring of long-horizon reproduction. Experiments on PaperBench Code-Dev show that DeepRepro consistently outperforms strong scientific and commercial code-agent baselines.
CommentsAccepted by CIKM2026 Demo Track