BVB:通过Blender中的程序化重建对智能体视频理解进行基准测试
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
浏览论文内容
中文总结 AI 辅助
提出BVB基准,要求智能体将真实视频程序化重建为Blender场景,以双VQA和潜在相似性评估视频理解,发现最佳模型视觉相似度高但事实保留不足。
中文摘要 AI 辅助
多模态智能体可以通过编码在Blender等软件中创建复杂视频,而无需依赖扩散模型。然而,视频理解基准仍然主要通过问答来评估模型。如果智能体真正理解视频,它可以通过编程方式重建视频。我们引入了BVB,即Blender-VideoBench,一个通过要求智能体将真实世界视频重建为动画Blender场景来测试这种能力的基准。为了确保公平比较,每个智能体通过一个轻量级框架Mini-BVB在相同的沙箱中、在共享的成本限制下编程重建。该基准从动画相机渲染每次重建,并从两个维度进行评估:(1)双VQA衡量重建保留了多时空事实。(2)潜在相似性衡量重建与源视频在感知上的匹配程度。我们的总分,即平方根均值,有利于均衡性能。我们评估了来自10个模型家族的51种配置,并分析了语义保留、感知相似性、推理努力和成本。最佳模型达到88.6的潜在相似性,但仅保留了53.7%的源正确时空答案。额外的推理改善了视觉相似性,但并未缩小事实准确性上的差距。在一项涉及15名评估者和五种配置的盲测研究中,潜在相似性与人类偏好强烈相关。这些结果表明,程序化重建是测试智能体视频理解的一种可行方法,而语义保留仍然是主要挑战。
英文摘要
Multimodal agents can create complex videos in software such as Blender by writing code instead of using diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconstruction through a lightweight harness, Mini-BVB, in an identical sandbox under a shared cost limit. The benchmark renders each reconstruction from its animated camera and evaluates it on two axes: (1) Dual VQA measures how many spatiotemporal facts the reconstruction preserves. (2) Latent Similarity measures how closely the reconstruction matches the source video perceptually. Our overall score, a square-root mean, favors balanced performance. We evaluate 51 configurations from 10 model families and analyze semantic retention, perceptual similarity, reasoning effort, and cost. The best model reaches 88.6 Latent Similarity but retains only 53.7% of the spatiotemporal facts from the source video. Additional reasoning improves perceptual similarity but does not close this gap. In a blind study with 15 raters and five configurations, Latent Similarity correlates strongly with human preference. These results show that programmatic reconstruction is a viable test of agentic video understanding, and that semantic retention remains the main challenge.
发表机构
- University of Rochester(罗切斯特大学)
- Sony Group Corporation(索尼集团公司)
- Carnegie Mellon University(卡内基梅隆大学)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。