AI 中文总结
本研究提出电路对齐分数(CAS),通过图核比较跨域类别特定电路,无需目标域数据即可预测OOD泛化,在PACS上秩相关达0.88,优于现有方法。
AI 中文摘要
能否仅从训练模型的权重预测分布外(OOD)泛化,而无需任何目标域数据?现有的表征相似性度量(CKA、SVCCA、RSA)比较激活而非预测泛化。我们证明它们对计算图中的结构重路由(即分布偏移所引发的变化)不敏感。我们通过电路对齐分数(CAS)填补了这一空白,该分数通过图核比较跨域的类别特定电路,并分解为同类一致性和跨类混淆。将CAS视为域分布上的勒贝格积分,我们证明其蒙特卡洛估计能恢复学习者按OOD准确率的真实排名,成对反转误差以$O(1/M)$速率消失,其中$M$为采样域数量。在PACS上的48个学习者中,CAS与OOD准确率的秩相关达到0.88,而CKA为0.58,SVCCA为0.23,RSA为0.14,在其他基准上以及甚至与依赖数据的方法相比也呈现类似趋势,使其成为首个可证明一致的分布鲁棒性预测器,且无需目标域数据或标签。代码可在以下网址获取:this https URL
英文摘要
Can out-of-distribution (OOD) generalization be predicted from a trained model's weights alone, without any target-domain data? Existing representational similarity metrics (CKA, SVCCA, RSA) compare activations rather than forecast generalization. We show they are provably insensitive to structural rerouting in the computational graph, the very change distribution shift induces. We close this gap with the Circuit Alignment Score (CAS), which compares class-specific circuits across domains via graph kernels, decomposed into same-class coherence and cross-class confusion. Casting CAS as a Lebesgue integral over the domain distribution, we prove its Monte Carlo estimate recovers the ground-truth ranking of learners by OOD accuracy, with pairwise inversion error vanishing at rate $O(1/M)$, where $M$ is the number of sampled domains. Across $48$ learners on PACS, CAS attains $0.88$ rank correlation with OOD accuracy, versus $0.58$ (CKA), $0.23$ (SVCCA), and $0.14$ (RSA), with similar trends on other benchmarks and even against data-dependent methods, making it the first provably consistent predictor of distributional robustness requiring neither target-domain data nor labels. The code is available at: https://github.com/ayanban011/ACE
CommentsAccepted at NeurIPS 2026