arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16198cs.CVcs.AI

选择合适的图像进行分类:远程皮肤病学中的可靠输入选择

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

  • University of Basel(巴塞尔大学)
  • Lucerne University of Applied Sciences and Arts(卢塞恩应用科学与艺术大学)
  • University Hospital Basel(巴塞尔大学医院)

机构由 AI 辅助整理,请以论文原文为准。

Fabian Gröger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly

AI总结:

该研究针对远程皮肤病学中模型因图像分布偏移分类错误的问题,提出可靠输入选择任务,测试多种无预训练数据的选择器,发现其均难缩小与神谕的差距,置信度融合马氏距离的选择器效果仍有限。

AI中文摘要:

皮肤病学模型在远程皮肤病学场景中面临分布偏移问题,提交的图像在光照、角度、距离、焦点和构图方面与训练数据存在差异。这些测试图像是普通临床照片,但部分超出模型的训练条件,由于训练与部署间采集方式的偏移,模型常对其分类错误。当同一病例存在多张图像(同一患者或同一病灶的多张照片)时,提升准确率的自然方式是选择模型最可能正确分类的图像,我们将此任务称为可靠输入选择。在六个皮肤病学数据集和九个冻结骨干网络上,若存在正确分类的图像,能为每个病例选择该图像的“神谕(oracle)”可使加权F1平均提升约20个百分点,该神谕是能查看标签的上限,而选择器必须盲目选择。在实践中实现这一增益难度较大。无需预训练数据的选择器可适用于任何冻结模型,包括那些数据不公开的模型,它必须从模型推理时暴露的量(嵌入、嵌入范数和置信度)判断可靠性。我们对四种此类无训练数据的选择器进行基准测试:嵌入范数、同一病例图像间的邻域共识、小扰动下预测的稳定性以及模型自身的置信度。无训练数据的选择器均未大幅缩小该神谕差距,其中最佳的是模型自身的置信度,但它在临床数据集上仅能恢复小部分差距;少量带标签的参考集也无帮助,整体最佳选择器(置信度与马氏距离的融合)仍留有大部分差距。据我们所知,这是首个引入并基准测试可靠输入选择的研究,该任务具有临床重要性且尚未解决。

英文摘要:

Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall outside the model's training conditions, leading the model to often misclassify them due to shifts in acquisition between training and deployment. When multiple images of the same case exist (several photos of one patient or lesion), a natural way to improve accuracy is therefore to select the image the model is most likely to classify correctly. We call this task reliable-input selection. An oracle that, for each case, selects a correctly classified image when one exists raises weighted F1 by about 20 percentage points on average across six dermatology datasets and nine frozen backbones. This oracle is an upper bound that sees the labels, whereas a selector must choose blindly. Capturing this gain in practice is hard. A selector that needs no pretraining data applies to any frozen model, including those whose data is not public. It must judge reliability from quantities the model exposes at inference: its embeddings, their norms, and its confidence. We benchmark four such training-data-free selectors: the embedding norm, the neighborhood consensus among a case's images, the stability of the prediction under small perturbations, and the model's own confidence. No training-data-free selector substantially narrows this oracle gap. The best of them is the model's own confidence, but it recovers only a small part of the gap on the clinical datasets. A small labeled reference set does not help either: the best selector overall, a fusion of confidence and Mahalanobis distance, still leaves most of the gap. To our knowledge, this is the first study to introduce and benchmark reliable input selection, a clinically important, unsolved task.

↑