AI 中文总结
研究探讨冻结基础编码器问题,引入CANDOR度量,其等规模库对称使机会水平固定为二分之一。经多编码器、多数据集和大量图像实验,校正改变结论,发现各编码器虽不盲目但都弱,CANDOR可提前标记支持不佳的发现。
AI 中文摘要
冻结编码器的选择依据是轻量级头部从其特征中读取发现的能力,而非几何结构是否能区分。最近邻不一致性虽能区分,但在不均衡库的情况下,相反标签的邻居在密度上占优,而非几何结构。我们引入了CANDOR,一种不一致性度量,其等规模库在标签交换下是对称的,将机会水平精确固定为二分之一。通过22个编码器、来自7个领域的20个数据集和605,443张图像的实验,这种校正改变了结论。几乎在所有地方崩溃率都低于机会水平,没有编码器是完全盲目但都很弱。最好的胸部模型读取气胸的AUROC为84.5,但仍有18.4%的阳性病例在同一家医院中与相反标签的胶片距离更近。同一编码器对鸟类物种的分辨率为4.5,对胸部发现的分辨率为42.8,对青光眼的分辨率为49.8,处于机会水平且比随机权重更差。这种情况限制了任何Lipschitz头部的归一化余量,但11个头部中有一个在除2.百分8%的情况外全部正确,而另一个头部错过35.9%:缺陷在于选择而非信息。擦除保留与崩溃相关;我们未检测到与发现的目标、规模、近期性或大小有关联。由于机会水平固定,CANDOR可在任何头部训练前读取,标记出冻结编码器支持不佳的发现。
英文摘要
A foundation encoder is pretrained once on a large image corpus and then reused with its weights frozen. Each new task is solved by training a small head on the features it produces. This setup is common in medical imaging, where labeled cases are scarce and a frozen encoder can be reused across findings. All downstream tasks then depend on the class separation present in that fixed feature space. Encoder selection usually uses the area under the receiver operating characteristic curve (AUROC) of a trained downstream head. AUROC measures the predictive information a head can extract from the features, but it does not measure class separation in the frozen feature space. A positive image that lies near a negative image in feature space has a bounded normalized margin under any Lipschitz head. Chance-calibrated neighborhood discordance (CANDOR) measures this feature-space separation without training a head. For a positive image, it compares the k nearest positive-label neighbors with the k nearest negative-label neighbors, among images acquired the same way. Discordance rate D is the share of positives whose opposite-label neighbors are nearer. Drawing the 2 candidate sets at equal size makes the labels interchangeable, so label-independent features have chance level D=1/2 without simulation. We apply CANDOR to 22 frozen encoders on 605,443 images from 20 public datasets, covering 8 binary tasks in 7 domains. On every task, the best encoder is below 1/2. A discordant image bounds the normalized margin of every Lipschitz head on that encoder. A selector that is shown the true label and chooses among 11 encoders reduces the miss rate from 0.359 to 0.028. Discordance is associated with occlusion retention and with none of pretraining objective, parameter count, release year, or finding size. The fixed chance level lets D be computed for a frozen encoder before any downstream head is trained.