TerraVis:基于MLLM工作流的文生图世界一致性评估
TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
浏览论文内容
中文总结 AI 辅助
针对文生图模型生成图像违反现实世界合理性的问题,提出TerraVis框架,通过结构化分类体系和MLLM多阶段评估量化世界一致性违规,在多个模型和基准上达到与人类判断最强相关性,补充了传统评估指标。
中文摘要 AI 辅助
近期的文生图模型在照片真实感、美学和文本-图像对齐方面取得了显著进展。然而,视觉上吸引人的图像仍可能违反现实世界的合理性,表现出畸形的物体结构、不可能的解剖结构、物理上不合理的交互或不一致的空间关系。现有的保真度、美学、偏好或对齐指标无法很好地捕捉此类失败。为解决这一差距,我们引入了TerraVis,一个用于评估生成图像中世界一致性(world-grounded visual consistency)的框架。TerraVis定义了一个结构化的世界一致性违规分类体系,涵盖物体级、交互级和场景级失败,并采用多阶段评估框架来识别和量化这些违规。给定一张图像,TerraVis首先使用MLLM评估其是否适合进行评估,然后检测18种分类学定义的违规类型,并将其分类为轻微或严重,以得出整体世界一致性得分。在两个广泛使用的基准测试中,针对多种开源和专有文生图模型,TerraVis在现有指标中实现了与人类对世界一致性判断的最强相关性。我们的基准测试结果进一步表明,在传统指标上表现优异的模型仍可能表现出大量的世界一致性失败。这些发现凸显了世界一致性作为一个互补的评估维度,并证明TerraVis能够系统性地量化、诊断和比较此类失败。我们的代码公开可用,网址为https://this https URL。
英文摘要
Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationships. Such failures are not well captured by existing fidelity, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations spanning object-, interaction-, and scene-level failures, and employs a multi-stage evaluation framework to identify and quantify them. Given an image, TerraVis first uses an MLLM to assess its eligibility for evaluation, then detects violations across 18 taxonomy-defined types and classifies them as minor or major to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis achieves the strongest correlation with human judgments of world consistency among existing metrics. Our benchmark results further show that models that achieve strong performance on conventional metrics can still exhibit substantial world-consistency failures. These findings highlight world consistency as a complementary evaluation dimension and demonstrate that TerraVis enables systematic quantification, diagnosis, and comparison of such failures. Our code is publicly available at https://github.com/ShyFoo/TerraVis.
发表机构
- AIML, Adelaide University(阿德莱德大学AIML)
- xAI
- Responsible AI Research Centre(负责任AI研究中心)
机构由 AI 辅助整理,请以论文原文为准。