arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LiveEvalBench:面向网页生成的开放世界评估

LiveEvalBench: Toward Open-World Evaluation for Web Generation

Yiyao Wang, Zhen Wen, Yinghao Tang, Yixiao Fu, Lin Yuan, Xiaolau Zhang, Jun Zhou, Wei Chen

arXiv 2608.03689首次发表:更新:

AI 中文总结

LiveEvalBench是将网页生成评估转为智能体驱动的自适应可扩展过程的自动化框架,经实验验证其与人类专家判断高度一致,可提供前沿模型网页生成能力的细粒度洞察

AI 中文摘要

大型语言模型在合成可执行前端项目方面的能力日益增强,但现有基准仍将网页生成视为静态评估问题。我们认为前端制品需要不同的评估范式:它们具有交互性而非静态性,存在多样但同等有效的实现方式,且演化速度快于刚性流水线的适配能力。为解决这些不足,我们提出LiveEvalBench,这是一个自动化框架,将网页生成评估重新定义为智能体驱动的、自适应且可扩展的过程。LiveEvalBench将评估实例化为协作评审工作流,其中构建工程师、代码工程师和UI测试人员在前端项目的全生命周期中共同收集证据,从部署、代码检查到基于浏览器的交互。为处理实现多样性,自适应协议结合了用于跨模型可比性的共享 rubrics(评分标准),以及针对每个制品定制的基于实现的标准。该框架还支持增量集成新的评估角色和评估维度,无需重新设计流水线。在不同的现实世界网页生成场景中进行的实验表明,LiveEvalBench与人类专家判断高度一致,并能提供前沿模型网页生成能力的细粒度洞察。代码可在this https URL获取

英文摘要

Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑