arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于最优传输的半监督域适应的共形风险最小化方法

Conformal Risk Minimization for Semi-Supervised Domain Adaptation via Optimal Transport

Manos Giannopoulos, Yi Shen, Michael M. Zavlanos

arXiv 2608.23153首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对半监督域适应场景中现有方法缺乏不确定性量化的问题,提出将共形风险最小化与最优传输结合的端到端框架,生成紧凑且覆盖有效的预测集。

AI 中文摘要

在高风险医疗应用中,机器学习模型常基于某一患者群体的数据进行训练,并部署到另一群体,由此产生的分布偏移会降低模型的准确性和可靠性。半监督域适应(SSDA)通过利用源域的有标签数据提升标签稀缺的目标域上的模型性能,然而现有SSDA方法主要针对点预测准确性优化,未提供原则性的不确定性量化——而这是获得临床信任的前提。共形预测(CP)可解决这一局限,它能提供具有严格无分布覆盖保证的预测集,但将CP事后应用于预训练模型会产生过大的预测集,因为SSDA预训练方法未考虑决定共形集大小的非一致性得分几何。共形风险最小化(CRM)已在全监督场景中通过将CP目标直接整合到模型训练来解决该问题,但它需要大量有标签数据集以在训练期间计算非一致性阈值,而这正是SSDA场景中稀缺的数据。我们提出一种将CRM整合到SSDA训练目标中的端到端框架,使CRM能在有限的目标域有标签数据场景中有效运行。核心思路是利用最优传输(OT)为无标签目标实例生成伪标签,为CRM提供额外的训练信号,使其仅使用少量目标域有标签集即可运行。这一方法得到的模型同时针对域不变性和共形效率进行优化,生成的预测集紧凑、覆盖有效,且支持特定域约束,例如在皮肤病变分类中排除相互矛盾的诊断结果。

英文摘要

In high-stakes healthcare applications, machine learning models are frequently trained on data from one patient population and deployed on another, creating a distribution shift that degrades both accuracy and reliability. Semi-Supervised Domain Adaptation (SSDA) addresses this by leveraging labeled data from some source domain to improve model performance on a target domain where labels are scarce. However, existing SSDA methods optimize primarily for point-prediction accuracy and offer no principled uncertainty quantification --- a prerequisite for clinical trust. Conformal Prediction (CP) can address this limitation by providing prediction sets with rigorous, distribution-free coverage guarantees. However, applying CP post-hoc to a pre-trained model can yield prohibitively large prediction sets, as SSDA pre-training methods do not account for the nonconformity score geometry that determines conformal set size. Conformal Risk Minimization (CRM) has been used to resolve this issue in the fully supervised setting by integrating the CP objective directly into model training, but it requires a large labeled dataset to compute nonconformity thresholds during training, precisely the data that is scarce in the SSDA regime. We propose an end-to-end framework that integrates CRM into the SSDA training objective, enabling effective CRM in the limited-labeled-target-data regime. The key idea is to utilize Optimal Transport (OT) to generate pseudolabels for unlabeled target instances, providing the additional training signal needed by CRM to operate using only a small labeled target set. This results in a model jointly optimized for domain invariance and conformal efficiency, producing prediction sets that are compact, coverage-valid, and support domain-specific constraints such as excluding mutually contradictory diagnoses in skin lesion classification.

Comments18 pages, 2 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑