发表机构
Fudan University; City University of Hong Kong; Chinese University of Hong Kong; Singapore Management University; HKUST (GZ); Monash University Malaysia; Harbin Institute of Technology(复旦大学; 香港城市大学; 香港中文大学; 新加坡管理大学; 香港科技大学(广州); 莫纳什大学马来西亚分校; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LoLBench是一个多语言基准,通过100个长周期提案任务评估编码智能体的感知与实现能力,发现最佳智能体仅解决14%任务,代码定位是主要瓶颈,提供文件树可提升16-22个百分点。
AI 中文摘要
现代编码智能体能够交付越来越大的仓库级变更,最近的基准测试通过强调具有大型参考实现的长周期任务来反映这一点。许多基准测试评估编码智能体的实现能力,即根据详细规格生成正确的代码编辑。然而,实际的模块化开发任务还需要感知能力,即将用户意图和高层设计落地为规格说明。我们引入LoLBench,通过在大规模软件系统上的整个提案到实现过程来评估这两种能力。它是一个多语言基准测试,包含5个领域的29个软件系统中的100个任务。每个任务提供一份人工撰写的增强提案,包含用户意图和高层设计。平均而言,提案包含约5000个单词,软件系统包含240万行源代码(LoC),实现拉取请求(PR)大约更改5500行代码。在我们评估的28个智能体中,最好的智能体仅解决了14%的任务,并实现了52.7%的失败到通过(F2P)通过率。失败分析确定不完整的代码定位是主要瓶颈,而提供参考派生的文件树以及API规格将解决率提高了16--22个百分点(2.4--17倍),最高达到34%。这些结果表明,感知和实现在大规模软件系统的实际模块化开发中仍然是编码智能体的核心挑战。LoLBench可在https URL上获取。
英文摘要
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.