arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03715cs.CVcs.AIcs.GR

4DCodeBench:动态场景逆向图形智能体基准测试

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu

首次发表
浏览论文内容

中文总结 AI 辅助

提出4DCodeBench基准,通过代码生成实现4D逆向图形,评估智能体从视频重建动态场景的能力,发现静态重建强但动态重建弱,为追踪进展提供测试平台。

中文摘要 AI 辅助

我们提出了4DCodeBench,一个通过代码生成进行4D逆向图形的基准测试,其中智能体将视频中的动态场景重建为可执行的图形程序。为了实现这一目标,智能体必须将视觉观察转化为场景结构和动态的紧凑表示,通过实现物理模拟等抽象来重现复杂行为。为了评估这一能力,我们策划了一组真实世界视频,并构建了涵盖多种物理现象(包括变形、流体流动和断裂)的合成场景。我们对前沿模型进行了广泛的基准测试,发现强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench为追踪智能体通过代码解释世界动态的进展提供了一个测试平台。我们的基准测试可在以下网址获取:此https URL

英文摘要

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench

发表机构

  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • Stanford University(斯坦福大学)
  • Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑