AI 中文总结
本文提出组合目标,通过有界观察者衡量图像部分间关系,改进谱epiplexity方法,在ImageNet上实现更强分类性能,并揭示谱多样性、可预测性与下游效用的区别。
AI 中文摘要
智能有多种定义方式。其中一种定义将智能视为对可学习新奇性的追求。然而,若无法将所学结构组合起来以采取行动并实现目标,可学习新奇性可能毫无意义。可学习新奇性建立在epiplexity之上,这是一种通过有界观察者来衡量数据中可学习结构的方法。在本文中,我们研究了一种用于计算epiplexity的闭式谱近似方法。我们采用固定迹约束,并发现epiplexity目标偏好将谱质量更均匀地分布,而非将其集中在少数方向上。然而,一种表示可能将信息分散到许多方向,却未将这些信息组织成对特定任务有用的特征。为解决这一差距,我们提出了一种组合目标,其观察者衡量图像各部分及交互之间的关系。我们仅使用掩蔽部分将其与原始谱目标进行比较。在我们的ImageNet训练运行中,具有掩蔽部分的谱目标产生了几乎最大程度分散的表示,同时在大多数评估中实现了最强的冻结特征分类性能,将线性探针准确率提高至整图基线的两倍以上。在多个图像基准上,改变观察者所见比添加关系和交互标记更为重要。同时,我们面向预测的组合目标产生了显著更好的保留观察者预测,但分类性能相对较弱,这表明谱多样性、可预测性和下游效用是不同的属性。这些结果表明,谱扩展的有用性不仅取决于保留了多少结构,还取决于观察者向目标提供了哪些关系。
英文摘要
Intelligence is defined in many ways. One of these definitions defines intelligence as the pursuit of learnable novelty. However, learnable novelty can be meaningless without the ability to compose the learned structures to take action and achieve goals. Learnable novelty builds on epiplexity, which is a way to measure learnable structure in data through a bounded observer. In this paper, we investigate a closed-form spectral approximation to compute epiplexity. We use a fixed-trace constraint and find that the epiplexity objective prefers a more uniform distribution of spectral mass rather than concentrating it in a small number of directions. However, a representation may spread information across many directions without organizing that information into features useful for a particular task. To address this gap, we propose a compositional objective whose observer measures the relationships between the parts and interactions of an image. We compare it with the original spectral objective given only the masked parts. In our ImageNet training runs, the spectral objective with masked parts produces an almost maximally spread representation while achieving the strongest frozen-feature classification performance on most evaluations, more than doubling the linear-probe accuracy of the whole-image baseline. Across multiple image benchmarks, changing what the observer sees matters more than adding relation and interaction tokens. At the same time, our prediction-oriented compositional objective produces substantially better held-out observer prediction but relatively weaker classification, revealing that spectral diversity, predictability, and downstream utility are distinct properties. These results suggest that the usefulness of spectral spreading depends not only on how much structure is preserved, but on which relationships the observer makes available to the objective.