预训练图像编码器中的表征风险
Representation Risk in Pretrained Image Encoders
查看机构详情
- Carleton University(卡尔顿大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出表征风险概念,证明预训练编码器选择显著影响下游预测,通过多任务比较提出基准测试、锁定验证选择的工作流程,并实现为LOOKAGAIN-ML软件包。
中文摘要 AI 辅助
应用研究人员越来越多地使用预训练编码器将图像转换为特征,然后在下游预测模型中使用这些特征。编码器通常被视为实现细节。我们表明,它反而可能是模型不确定性的一个重要来源。我们将这种不确定性称为表征风险:合理的预训练编码器将相同的图像映射到不同的特征空间,并可能从预测性能中产生截然不同的样本外结论。我们在涉及房价、赛马表现、乳腺癌组织学、胸部X光片、连续面部年龄和水稻病害的应用中比较了十个现代和传统冻结编码器。在常见的维度控制、头部和组安全分割下,验证选择SigLIP 2用于房价,将测试R²从ResNet50的0.396提高到0.629,选择DINOv2用于赛马,将R²从0.029提高到0.105。没有编码器在所有任务中都是最好的。候选程序使用训练数据构建,并在单独的验证分区上进行比较。所选程序在房价上达到0.658,在肺炎上达到0.979的准确率。固定分割的增益在赛马和水稻上较小,而重复分区揭示了赛马特征并集的不稳定性。连续年龄选择SigLIP 2,平均绝对误差为4.786年。主要的表征差距在神经头部、类似大小的DINOv2和ViT模型以及有限适应中持续存在。这些结果支持一个简单的工作流程:对合理的表征进行基准测试,在锁定的验证数据上进行选择,仅在单独的验证证据证明额外成本合理时才进行组合,并报告配对和分割级别的不确定性。我们在LOOKAGAIN-ML中实现了这一工作流程,该软件包用于进行本文的分析。
英文摘要
Applied researchers increasingly convert images into features with pretrained encoders, then use those features in a downstream prediction model. The encoder is often treated as an implementation detail. We show that it can instead be a consequential source of model uncertainty. We call this uncertainty representation risk: plausible pretrained encoders map the same images into different feature spaces and can yield sharply different out-of-sample conclusions from predictive performance. We compare ten modern and legacy frozen encoders across applications involving house prices, racehorse performance, breast-cancer histology, chest radiographs, continuous facial age, and rice disease. With common dimension control, heads, and group-safe splits, validation selects SigLIP 2 for houses, raising test $R^2$ from 0.396 for ResNet50 to 0.629, and DINOv2 for horses, raising $R^2$ from 0.029 to 0.105. No encoder is best in every task. Candidate procedures are constructed using training data and compared on a separate validation partition. The selected procedure reaches 0.658 for houses and 0.979 accuracy for pneumonia. Fixed-split gains are small for horses and rice, while repeated partitions reveal instability in horse feature union. Continuous age selects SigLIP 2 at 4.786 years MAE. The principal representation gaps persist with neural heads, similarly sized DINOv2 and ViT models, and limited adaptation. These results support a simple workflow: benchmark plausible representations, select on locked validation data, combine only when separate validation evidence justifies the additional cost, and report paired and split-level uncertainty. We implement this workflow in LOOKAGAIN-ML, the software package used to conduct the analyses in this paper.