发表机构
Institut Polytechnique de Paris; École Polytechnique; Télécom Paris(巴黎综合理工学院; 巴黎综合理工学院; 巴黎高等电信学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出3DHarnessBench基准,通过四种递增交互设置评估前沿VLM从输入恢复3D几何并生成Blender代码的智能体能力,发现函数调用访问显著提升性能但模型间差异大。
AI 中文摘要
我们引入了3DHarnessBench,这是一个基准测试,用于评估前沿视觉语言模型(VLMs)的智能体能力,即从各种输入中恢复3D几何形状并将其转换为Blender Python代码。与之前使用固定输入(例如,单一渲染或文本描述)提示VLM的框架不同,3DHarnessBench评估了四种不同的harness设置,这些设置逐步启用主动的智能体探索,并借助最近的Blender MCP功能实现。我们的层级从单视图、多视图、主动视觉(任意视点访问)到全3D交互(通过Blender函数调用完全访问目标对象),探测了模型在视觉感知和主动推理、工具调用以及自我纠正方面的能力。我们观察到,所有前沿模型恢复3D几何形状的能力随着更丰富的函数调用访问而显著提高,尽管这些改进强烈依赖于模型,揭示了高度不均衡的智能体3D到代码能力。我们将发布基准测试、代码、输出和智能体轨迹,以实现可复现的3D评估。
英文摘要
We introduce 3DHarnessBench, a benchmark that evaluates the agentic ability of frontier vision-language models (VLMs) to recover 3D geometry as Blender Python code from a variety of inputs. Unlike previous frameworks that prompt the VLMs with a fixed input (e.g., a single rendering or a text description), 3DHarnessBench evaluates four separate harness settings that progressively enable active agentic exploration, facilitated by recent Blender MCP functionality. Our hierarchy from Single-view, Multi-view, Active Visual (arbitrary viewpoint access), and Full 3D Interaction (complete access to the target object through Blender function calls) probes the models' abilities in both visual perception and active inference, tool calling, and self-correction. We observe that the ability of all frontier models to recover 3D geometry improves significantly with richer function call access, although the improvements are strongly model-dependent, revealing highly uneven agentic 3D-to-code capabilities. We will release the benchmark, code, outputs, and agent trajectories for reproducible 3D evaluation.