AI 中文总结
本文提出WorldBench基准测试与评估器,针对大语言模型生成的Three.js体素世界,解决单一视角评估不可靠问题,通过代码与视觉双维度评估,降低特征遗漏的误判,评估了五个前沿模型。
AI 中文摘要
大型语言模型如今可编写完整且具交互性的3D世界代码,但自动评估这些世界的可靠性不足。现有评估方法仅从单一视角判断输出:视觉语言模型对少量渲染快照打分,或语言模型读取源代码。针对五个前沿模型生成的世界,研究发现两种视角在32%的必要项上存在分歧,分歧多源于代码未在任何帧中显示,且固定视角会遗漏近距离细节。本文提出WorldBench,这是一个针对大语言模型生成的开放式Three.js世界的基准测试与评估器。从一个描述包含十个生物群系、物理系统及昼夜与季节循环的漂浮体素岛的提示词出发,该评估器会探索运行中的世界,控制其时钟、环绕世界,并派遣导航智能体对每个生物群系取景,同时读取其所见内容的代码。两种视角均不可单独信赖:代码引用仅当为源代码中存在的文本时才有效,视觉主张则需对照可测量属性的像素进行核查。一项通过刻意移除特征的变异测试显示,仅基于代码的评估会对五个被移除特征中的四个给予满分,因为这些特征的代码仍保留在文件中。本文提出的评估器将被移除特征的保留分数降低了三分之一(从7.11分中的5.44分降至3.55分),且其仍认可的内容大多是存在但从未运行的代码。本文评估了五个前沿模型:Claude Fable 5.1、GPT-6 Astra、Kimi K3、Grok 4.7和Gemini 3.1 Pro。代码、提示词、测试及评估器配置可在指定网址获取。
英文摘要
Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents. We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds. From one prompt describing a floating voxel island with ten biomes, physics, and day/night and seasonal cycles, the judge explores the running world, controlling its clock, orbiting it, and sending a navigator agent to frame each biome, and reads the code for what it sees. Neither channel is trusted on its own: a code quote counts only if it is text the source contains, and visual claims are checked against measured pixels where the property is measurable. A mutation test, in which we remove features by construction, shows that code-only judging gives full credit to four of five removed features, because their code remains in the file. Our judge cuts the points kept on removed features by a third (5.44 to 3.55 of 7.11), and what it still credits is mostly code that exists but never runs. We evaluate five frontier models: Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7 and Gemini 3.1 Pro. Code, prompt, tests and judge configuration are available at https://github.com/KrishBakshi/worldbench