位移几何捕捉跨模型与跨模态的柏拉图式共享现实
Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities
浏览论文内容
中文总结 AI 辅助
本文发现模型间共享的是表征空间中的位移方向而非绝对位置,提出影子投射方法,通过缓存前向传播实现无需微调的能力迁移,并以SHADOWCLIP验证其高效性。
中文摘要 AI 辅助
柏拉图式表征假说(PRH)声称,独立训练的模型会收敛于一个共享的关于现实的统计模型,然而近期研究发现模型之间仅存在微弱的点级相似性。在本文中,我们表明模型之间共享的并非样本在表征空间中的位置,而是它们之间的方向(位移向量)。在单一正交对齐(仅旋转和反射)下,这些位移向量在44个独立训练的视觉和语言编码器中得到了显著保留,这些编码器跨越了不同模态和不对称能力组合,与PRH的证据一致。而样本的绝对位置则并非如此,这与近期的反证一致。两者均源于一个统一的分解:表征被分为一个跨模型线性对齐的共享语义成分和一个不对齐的私有能力成分。我们将这种几何追溯到概念级结构:在模型内部,父概念与其子变体向量正交;在模型之间,概念位移是平行的。我们的理论可证伪地预测(且实验证实)微调保留点级相似性但破坏位移,而关系蒸馏则相反。一个重要的启示是,由于语义线性对齐而能力不对齐,能力可以通过一次缓存的源模型前向传播从一个模型导入到另一个模型。我们称之为“影子投射”。作为概念验证,我们的SHADOWCLIP实例以数量级更少的计算超越了强微调基线。缓存可以与开放模型权重一起发布,使得一个模型的能力可以被下载并导入到任意数量的其他模型中,而无需微调。
英文摘要
The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between models. In this paper, we show that what models share is not the location of samples in representation space, but the directions (displacement vectors) between them. Under a single orthogonal alignment--rotation and reflection only--these displacement vectors are substantially preserved across 44 independently trained vision and language encoders spanning modalities and asymmetric capability pairs, consistent with the PRH evidence. The samples' absolute positions are not, consistent with recent counter-evidence. Both arise from a single decomposition: representations split into a shared semantic component that is linearly aligned across models, and a private capability component that is not. We trace this geometry to concept-level structure: within a model, parent concepts are orthogonal to their child variation vectors; across models, concept displacements are parallel. Our theory falsifiably predicts (and experiments confirm) that fine-tuning preserves pointwise similarity but collapses displacement, and that relational distillation does the opposite. A major implication is that, because semantics align linearly but capabilities do not, capabilities can be imported from one model to another using a single cached forward pass through the source. We call this Shadow Casting. As a proof of concept, our SHADOWCLIP instantiation outperforms strong fine-tuned baselines at orders of magnitude less compute. A cache can be released alongside open model weights, letting one model's capabilities be downloaded and imported into any number of other models without fine-tuning.
发表机构
- Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。