arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时间-事件模型何时是浪费时间?在右删失下连接混合治愈模型与正-未标记学习用于二分类

When are time-to-event models a waste of time? Bridging mixture cure models and positive-unlabeled learning for binary classification under right-censoring

Qin Weng, Matthew M. Engelhard

arXiv 2609.19370首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文分析混合治愈模型与正-未标记学习在右删失二分类中的关系,发现仅在选择性删失且删失特征不可识别时MCM更优,否则PU学习性能相当且更稳健。

AI 中文摘要

在临床环境中,预测二分类结果常常因右删失而变得复杂,这使我们无法区分事件从未发生的个体(即阴性,或不易感者)与事件在删失时间之后发生的个体(即阳性,或易感者)。为了区分这些亚群,我们可以使用混合治愈模型(MCMs),该模型同时建模无限时间范围内的结局概率以及观察到事件时的其时间-事件(TTE)分布。然而,TTE方法需要事件时间戳,这些时间戳可能不可靠、获取成本高,或因临床护理差异而产生偏倚。或者,该任务可以表述为正-未标记(PU)学习,其中未观察到事件的个体被归为“未标记”而非阴性,认识到该组包括事件在删失后发生的阳性个体。在此,我们分析这些框架之间的关系,并为在它们之间进行选择提供基于证据的建议。我们首先证明,两个家族优化共享的似然,但不同之处在于它们如何约束易感患者在删失前事件发生的概率,这在PU学习中被称为标记倾向;并且MCM可以被视为标记倾向被参数化为TTE分布的PU模型,这使得模型可识别。在系统模拟和两个真实世界临床队列中,我们考察了包含TTE组件是否以及何时有利于二分类。我们表明,MCM仅在一种特定情况下被优先选择:在选择性删失下,当控制删失的确切特征无法被识别时。否则,PU学习实现同等性能,同时避免与TTE方法相关的挑战和潜在偏倚。

英文摘要

In clinical settings, predicting binary outcomes is often complicated by right-censoring, which prevents us from distinguishing individuals in whom the event never occurs (i.e., negative, or non-susceptible) from those where it occurs after the censoring time (i.e., positive, or susceptible). To discriminate between these subpopulations, we can use mixture cure models (MCMs), which model both the outcome probability over an infinite horizon and its time-to-event (TTE) distribution when observed. However, the TTE approach requires event timestamps, which can be unreliable, costly to obtain, or biased due to disparities in clinical care. Alternatively, the task can be formulated as positive-unlabeled (PU) learning, in which individuals without an observed event are grouped as "unlabeled" rather than negative, recognizing that this group includes positives whose event occurred after censoring. Here we analyze the relationship between these frameworks and provide evidence-based recommendations for choosing between them. We begin by showing that both families optimize a shared likelihood but differ in how they constrain the probability of event occurrence prior to censoring among susceptible patients, which in PU learning is known as the labeling propensity; and that an MCM may be viewed as a PU model whose labeling propensity is parameterized as a TTE distribution, which makes the model identifiable. In systematic simulations and two real-world clinical cohorts, we examine whether and when including the TTE component benefits binary classification. We show that MCMs are preferred in only one specific regime: under selective censoring where the exact features governing censoring cannot be identified. Otherwise, PU learning achieves equivalent performance while avoiding challenges and potential biases associated with the TTE approach.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑