ProgramDistill:从交互式Web应用到可验证的参考引导软件工程任务
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
- KAIST(韩国科学技术院)
- Microsoft Research Montréal(微软蒙特利尔研究院)
- Microsoft AI(微软人工智能)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ProgramDistill通过交互式参考应用生成可验证的软件工程任务,构建了包含4,063个任务的基准,评估编码智能体,发现GPT-6 Astra和Claude Opus 5在完整重建中成功率分别为49.2%和28.8%,并随恢复深度增加而下降。
AI中文摘要:
编码智能体通常通过问题或指令所指定的期望行为来评估。然而,在实际的Web开发中,智能体可能需要从可工作的软件中推断行为,并在不完整的应用程序中实现该行为。我们引入了ProgramDistill,一个基准测试,用于评估编码智能体在通过与功能完整的参考应用程序交互而发现的特性上的表现。我们通过将应用程序分解为不同粒度的特性来构建ProgramDistill,每个特性都关联了可通过其黄金补丁执行的可重放行为。我们的流水线mine-craft-patch在26个应用程序中发现了1,975个可重放验证的行为,并在无需人工干预的情况下构建了4,063个任务。在九个前沿编码智能体中,GPT-6 Astra和Claude Opus 5在完整应用程序重建的累积工作流中分别取得了49.2%和28.8%的成功率。在部分应用程序重建中,随着恢复深度从1增加到8,成功率从100%下降到64.0%,从96%下降到32%。因此,ProgramDistill提供了一个具有可控难度的可扩展基准,用于评估和诊断编码智能体,并为未来基于课程的学习提供了自然基础。
英文摘要:
Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications. We build ProgramDistill by factorizing applications into features of different granularities, each associated with replayable behaviors executable via its gold patch. Our pipeline, mine-craft-patch, discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks without human intervention. Across nine frontier coding agents, GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction. In partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8. ProgramDistill thus provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.