AI 中文总结
该研究针对冻结视觉编码器的操作洗白问题,提出SO-OPF读出方法,在Shapes3D、COCO等数据集上提升了因子分配准确率,同时暴露了渲染器特定的失效边界。
AI 中文摘要
冻结视觉编码器的组合分析应同时确定变化的内容与变化的位置。标准因子探针分别对这些轴进行评分,但可能会奖励重复使用同一预测槽的多种操作,我们将这种失效称为“操作洗白”。我们在支持×操作网格上引入了单单元留一法协议,以及SO-OPF(一种将单元能量分解为支持显著性和竞争性操作后验的读出方法)。该公式区分了聚合评分混淆的两个问题:当网格已知时,载体是否能组合保留的绑定;以及是否可从扁平单元标签中恢复该网格。在冻结DINOv3特征下,已知因子分配在Shapes3D-Extended上达到0.874的单射准确率,在全局图像不相交的COCO上达到0.799;从扁平标签学习分配时,分别达到0.769和0.762。在Shapes3D上的匹配轴感知监督下,因子载体相比密集载体将学习分配准确率从0.653提升至0.841,并消除了其洗白差距。SigLIP2复现了COCO上的分离结果。重建的MuJoCo基底暴露了一个边界:使用DINOv3时学习分配准确率为0.569,使用SigLIP2时为0.484,且存在显著的槽崩溃。因此,因子读出和单射评估在两个基底上恢复了保留的绑定,同时暴露而非隐藏了渲染器特定的失效边界,但并未确立从扁平标签的通用恢复能力。
英文摘要
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot. We call this failure operation laundering. We introduce an injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior. This formulation separates two questions that aggregate scores conflate: whether the carrier composes held-out bindings when the grid is known, and whether that grid can be recovered from flat cell labels. With frozen DINOv3 features, known factorial assignment reaches 0.874 injective accuracy on Shapes3D-Extended and 0.799 on globally image-disjoint COCO; learning the assignment from flat labels reaches 0.769 and 0.762, respectively. Under matched-axis-aware supervision on Shapes3D, the factored carrier improves learned-assignment accuracy from 0.653 to 0.841 over a dense carrier and eliminates its laundering gap. SigLIP2 replicates the COCO separation. A rebuilt MuJoCo substrate exposes a boundary: learned-assignment accuracy is 0.569 with DINOv3 and 0.484 with SigLIP2, with substantial slot collapse. Thus factored readout and injective evaluation recover held-out bindings on two substrates while exposing, rather than hiding, a renderer-specific failure boundary; they do not establish universal recovery from flat labels.