VideoVIBE:用于一次性交互式网站生成的基于视频的诊断基准
VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation
AI总结:
针对现有一次性交互式网站生成质量评估的不足,提出VideoVIBE基准与V2Lens多智能体系统,实验显示V2Lens可提升Video MLLM的评估性能。
AI中文摘要:
自然语言驱动的“ vibe coding”可实现视觉丰富且具交互性的网页应用的一次性生成,但对其质量的可靠评估却未能同步跟进。现有评估通常仅对孤立的产物或最终任务结果打分,仅能提供有限的证据说明出现了哪些故障以及为何发生。我们提出VideoVIBE,这是一种基于视频的基准,可将人类操作网页的记录转化为细粒度诊断任务。它包含约1700个诊断视频问答实例,这些实例源自生成网页中的6338个已验证故障,涵盖语义逻辑、视觉运动、结构时间及功能故障。诊断主要基于记录的呈现内容与行为,网页源代码作为补充上下文。我们进一步提出V2Lens,这是一种无训练、基于证据的多智能体系统,通过针对性的视觉及源代码验证,对基于视频的初始诊断提出质疑并进行选择性优化。在13种闭源及开放权重视频多模态大语言模型(Video MLLM)中,Gemini-2.5-Flash是表现最强的单模型,得分64.54,而V2Lens的得分达到71.72,提升了7.18个百分点。综合结果表明,基于视频的评估能够超越孤立产物与聚合结果,转向对生成应用质量的行为忠实且诊断信息丰富的评估。
英文摘要:
Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace. Existing evaluations often score isolated artifacts or final task outcomes, offering limited evidence about which failures occur and why. We introduce VideoVIBE, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks. It contains approximately 1.7K diagnostic Video QA instances derived from 6,338 verified failures across generated webpages, spanning semantic-logical, visual-motion, structural-temporal, and functional failures. Diagnoses are grounded primarily in recorded presentation and behavior, with webpage source code used as complementary context. We further propose V2Lens, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-code verification. Across thirteen closed-source and open-weight Video MLLMs, Gemini-2.5-Flash is the strongest standalone model with a score of 64.54, while V2Lens reaches 71.72, an improvement of 7.18 points. Together, our results show that video-grounded evaluation can move beyond isolated artifacts and aggregate outcomes toward a behaviorally faithful and diagnostically informative account of generated application quality.