arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ICAE-Bench:将编码智能体评估为交互式项目构建者

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

Zhongyuan Peng, Dan Huang, Chuyu Zhang, Caijun Xu, Changyi Xiao, Shibo Hong, David Lo, Lin Qiu, Xuezhi Cao, Jiyuan He, Yixin Cao

arXiv 2607.21217首次发表:更新:

发表机构

Meituan(美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对编码智能体在交互式项目构建中作用转变,现有基准未跟上的问题,提出ICAE-Bench基准。从模糊需求出发,经自动化用户智能体模拟动态范式,引入避免需求模糊性、确保用户模拟质量、公平评估开放式仓库的关键设计。

AI 中文摘要

近期出现的氛围编码工作流程正在改变对编码智能体的预期。智能体不再仅仅是在完全指定的指令下完成代码,而是越来越需要通过结合规划、需求澄清、工具使用、调试和仓库级构建等各种能力,将不完整的产品意图转化为可运行的软件。然而,现有基准尚未完全跟上这一转变,仍在静态、完全指定的任务上评估智能体。本文介绍了ICAE-Bench,这是一个用于在交互式项目构建设置下评估编码智能体的基准。基本思路是从模糊的产品需求开始,用自动化用户智能体模拟动态范式。为使该设置既现实又可评估,ICAE-Bench引入了三个关键设计。首先,为避免无约束模糊需求的模糊性,每个任务的模糊性源自具有可执行行为的精确真实开源仓库。其次,为确保高质量和可重复的用户模拟,ICAE-Bench通过用户智能体数据进行交互,让用户智能体揭示隐藏约束而不发明新需求或泄露实现工件。第三,为公平评估开放式仓库,ICAE-Bench使用标准化黑盒测试以及多维诊断,包括功能正确性、语义和API相似性、结构保真度、设计质量和交互质量。

英文摘要

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirement clarification, tool use, debugging, and repository-level construction. Yet existing benchmarks have not fully caught up with this shift, evaluating agents on static, fully specified tasks. In this paper, we introduce ICAE-Bench, a benchmark for evaluating coding agents under interactive project-building settings. The basic idea is to start from a fuzzy product requirement, simulating the dynamic paradigm with an automated User Agent. To make this setting both realistic and evaluable, ICAE-Bench introduces three key designs. First, to avoid the ambiguity of unconstrained fuzzy requirements, each task derives ambiguity from a precise real open-source repository with executable behavior. Second, to ensure high-quality and reproducible user simulation, ICAE-Bench grounds interaction through User Agent Data, allowing the User Agent to reveal hidden constraints without inventing new requirements or leaking implementation artifacts. Third, to evaluate open-ended repositories fairly, ICAE-Bench uses standardized black-box tests together with multi-dimensional diagnostics, including functional correctness, semantic and API similarity, structural fidelity, design quality, and interaction quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑