IWC-Bench:从软件测试视角评估Web应用生成
WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective
- Tencent(腾讯)
- Tsinghua University(清华大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
IWC-Bench从软件测试视角,通过代码覆盖率引导探索和状态转换图评估,在三个维度上评测LLM生成的Web应用,包含369个需求和5088个验收标准,与人类偏好一致性达85.3%。
AI中文摘要:
人工评估提供了对LLM生成的Web应用质量的直接度量。然而,通过自动化评估来拟合人类判断仍然具有挑战性。静态基准可以认可源代码中存在但在运行时不可达的功能。交互式基准会运行应用程序,但不完整的探索可能导致它们遗漏已实现的功能,并将应用程序缺陷与智能体执行失败混为一谈。为了解决这些局限性,我们提出了IWC-Bench,一个从软件测试视角评估Web应用生成的交互式基准。IWC-Bench对每个生成的应用程序进行插桩,并使用代码覆盖率来引导智能体通过用户模拟交互探索其功能。然后,它将交互轨迹抽象为状态转换图,并从三个维度评估应用程序:视觉美感、可用性和需求对齐。通过将探索与评分分离,IWC-Bench收集运行时证据,而不将探索限制在预定义的验收标准内。IWC-Bench包含369个真实世界用户需求和5,088个验收标准。对16个前沿LLM的评估揭示了三个维度上的不同优势,没有模型在每个维度上都领先。在从内部竞技场采样的197个验证会话中,IWC-Bench与人类偏好的一致性达到85.3%,且随着配对应用之间分数差异的增大,一致性通常会增加。进一步的实验表明,覆盖率引导提高了探索覆盖率,并且当替换评判模型时,模型排名保持稳定。
英文摘要:
Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose WebCraftBench, an interactive benchmark for evaluating web application generation from a software testing perspective. WebCraftBench instruments each generated application and uses code coverage to guide an agent in exploring its functionality through user-simulated interactions. It then abstracts the interaction trace into a state-transition graph and evaluates the application along three dimensions: visual aesthetics, usability, and requirement alignment. By separating exploration from scoring, WebCraftBench collects runtime evidence without constraining exploration to predefined acceptance criteria. WebCraftBench comprises 369 real-world user requirements and 5,088 acceptance criteria. Evaluation of 17 frontier LLMs reveals distinct strengths across the three dimensions, with no model leading on every dimension. On 197 validated sessions sampled from an internal arena, WebCraftBench achieves 85.3\% agreement with human preferences, with agreement generally increasing as the score difference between paired applications grows. Further experiments show that coverage guidance improves exploration coverage and the model rankings remain stable when the judge model is replaced.