arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EditCLEVR:用于以对象为中心表示的组合忠实性的配对场景干预基准测试

EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations

Anuraag Gadehothur Karnam, Tarunesh Sathish

arXiv 2607.22705首次发表:更新:

AI 中文总结

研究以对象为中心表示的组合忠实性,引入EditCLEVR基准测试,含配对场景及多种诊断与度量,如SGIA等,通过多种模型基线评估发现现有问题,能分别评估代码空间移动和解码对象属性正确性。

AI 中文摘要

以对象为中心的学习旨在将场景表示为其属性可在新组合中重用的对象。现有评估通常对分割、单图像因子预测或下游准确性进行评分,但这些测试并未直接询问在受控语义编辑下每个对象表示的行为是否正确。我们引入了EditCLEVR,这是一个配对场景干预基准测试,其中每个示例包含一对具有相同对象索引和场景布局的CLEVR风格渲染图,要么在一个已知对象上恰好有一个已知属性变化,要么是用于漂移测量的无编辑重新渲染。该协议包括用于表示变化定位和稳定性的无探针诊断,以及探针解码的语义忠实性度量,用于测试预测的场景变化是否与分布内和组合分布外(OOD)套件中的预期干预相匹配,从而能够分别评估代码空间移动和解码对象属性的正确性。我们引入了语义度量场景图干预准确性(SGIA),它要求完整的后场景预测正确,并且唯一预测的前后语义变化是预期的对象因子编辑。我们还建立了Delta-SGIA作为一个配套诊断,用于检查单站点变化模式,而无需完整的后场景图正确。对真实掩码主干、学习插槽模型、SAM 2 + 冻结ViT模型和一种掩码特征混合模型的基线评估表明,在真实实例掩码下,CoGenT-OOD核心退化可能会持续存在,掩码源占原生性能的一部分但不是全部,并且仅局部性或稳定性可能会夸大语义忠实性。代码可在此https URL获取。

英文摘要

Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations. Existing evaluations usually score segmentation, single-image factor prediction, or downstream accuracy, but these tests do not directly ask whether a per-object representation behaves correctly under a controlled semantic edit. We introduce EditCLEVR, a paired-scene intervention benchmark in which each example contains a before/after pair of CLEVR-style renders with the same object indices and scene layout, and either exactly one known attribute change on one known object or a no-edit re-render for drift measurement. The protocol includes probe-free diagnostics for representation-change localization and stability, together with probe-decoded semantic faithfulness metrics that test whether the predicted scene change matches the intended intervention across in-distribution and compositional out-of-distribution (OOD) suites, allowing code-space movement and decoded object-attribute correctness to be evaluated separately. We introduce the semantic metric Scene-Graph Intervention Accuracy (SGIA), which requires the full after-scene prediction to be correct and the only predicted before-to-after semantic change to be the intended object-factor edit. We also establish Delta-SGIA as a companion diagnostic that checks the single-site change pattern without requiring the full after-scene graph to be correct. Baseline evaluations on ground-truth-mask backbones, learned-slot models, SAM 2 + frozen-ViT models, and one mask-feature hybrid indicate that CoGenT-OOD-core degradation can persist under ground-truth instance masks, that mask source accounts for part but not all of native performance, and that locality or stability alone can overstate semantic faithfulness. Code is available at https://github.com/torux-bughunter/EditCLEVR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑