arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30647cs.CV

条件预测充分统计量用于视觉表征学习

Conditional Predictive Sufficient Statistics for Visual Representation Learning

Yuzhou Hong

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出条件预测充分统计量(CPSS)形式化视觉表征学习,通过余弦损失预测下一补丁嵌入作为代理,并指出停止梯度仅阻断常数解,真正充分统计量位于中间块,实验验证了其几何性质。

中文摘要 AI 辅助

有用的视觉表征是对已观察过去的统计量,它保留与未来共享的潜在因子,并丢弃补丁私有的噪声。我们将这一需求形式化为条件预测充分统计量(CPSS)。在图像补丁的共享因子模型下,过去与下一个补丁之间的互信息等于过去关于共享因子所携带的信息,减去下一个补丁本身无法揭示的余项。使用余弦损失预测下一个补丁嵌入,是对该嵌入方向的von Mises-Fisher模型的最大似然估计,因此是预测信息的可处理代理。相同的总体损失也可被常数嵌入最小化,因此停止梯度本身并不选择充分统计量;它仅阻止一步实现常数解的对称梯度。回归目标是一个浅层嵌入,这迫使网络输出回到该浅层范围,并将充分统计量留在中间块中。在MNIST和CIFAR-10上的小型因果Transformer用作诊断工具,而非排行榜。在MNIST上,未来偏移和停止梯度使探针准确率移动数十个百分点,CPSS读出峰值出现在输出之前。在CIFAR-10上,在相同的短预算且无增强的情况下,每个目标都接近像素上的线性分类器。与推导仍然匹配的是几何性质:CPSS输出是比其最佳中间块更差的读出,下一像素回归不承担该惩罚,且移除停止梯度会崩溃嵌入的有效秩,即使前置任务损失看起来完美。

英文摘要

A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.

发表机构

  • Zhejiang Sci-Tech University(浙江理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑