arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30102cs.LG

SMOTE-VAR:一种用于预测大学生抑郁缓解情况的不确定性感知过采样方法

SMOTE-VAR: An Uncertainty-Aware Oversampling Method for Predicting Depression Remission in University Students

Dang Nguyen, Arun Kumar A, Taylor A. Braund, Wu Yi Zheng, Debopriyo Bal, Leonard Hoon, Jill Newby, Helen Christensen, Svetha Venkatesh, Alexis Whitton, Sunil Gupta

首次发表
浏览论文内容

中文总结 AI 辅助

针对大学生抑郁预测中ML模型的类别不平衡问题,本文提出SMOTE-VAR过采样方法,通过高斯过程方差函数估计样本不确定性减少假阳性,在抑郁数据集上表现优于现有方法,助力优化精神健康护理。

中文摘要 AI 辅助

大学生患抑郁等常见精神健康问题的比例过高,这类问题会损害其学习、社会功能及整体健康。尽管正念、体育活动等生活方式干预可减轻症状,但许多学生无法达到症状缓解。开发识别预后不佳学生的新方法,可实现更早、更有针对性的干预。机器学习(ML)方法已越来越多地用于预测抑郁患者的缓解情况。然而,这些ML模型常面临类别不平衡问题,即缓解组与未缓解组的人数比例不均,这种不平衡会降低模型准确率并导致预测偏差。为解决该问题,研究常用流行的过采样策略SMOTE,但SMOTE存在显著缺陷:可能生成无效的合成少数类样本。在临床场景中,这些假阳性样本会导致错误的风险分层,可能延误对不可能缓解的患者升级必要护理。本文提出一种新型有效过采样方法以解决该缺陷,该方法利用高斯过程的方差函数估计生成的少数类样本的不确定性,从而减少假阳性。我们在收集的大学生抑郁数据集上验证了该方法,证明其在预测缓解情况(即治疗结局)方面优于现有过采样方法。通过改进对无应答者的可靠识别,该方法提供了一种稳健的计算工具,帮助临床医生快速转向辅助治疗,从而个性化并优化精神健康护理路径。

英文摘要

University students experience disproportionately high rates of common mental health conditions, such as depression, which can impair learning, social functioning, and overall well-being. Although lifestyle interventions such as mindfulness and physical activity can reduce the symptoms, many do not achieve symptomatic remission. Developing new approaches to identify students with poor outcomes could enable earlier and more targeted intervention. Machine learning (ML) methods have increasingly been used to predict remission in depressive patients. However, these ML models often suffer from class imbalance, where there may be an unequal proportion of people in the remitted group relative to the non-remitted group. This imbalance can reduce model accuracy and bias predictions. To address this, studies commonly employ the popular oversampling strategy SMOTE. However, SMOTE has a notable limitation: it may generate invalid synthetic minority samples. In a clinical context, these false positives can lead to incorrect risk stratification, potentially delaying necessary escalated care for patients unlikely to remit. In this paper, we introduce a novel and effective oversampling method that addresses this shortcoming. Our approach leverages the variance function of a Gaussian process to estimate the uncertainty of generated minority samples to reduce false positives. We validate our method on a depression dataset collected from university students and demonstrate that it is better than existing oversampling approaches in predicting remission (i.e., treatment outcome). By improving the reliable identification of non-responders, our method provides a robust computational tool to help clinicians rapidly pivot to adjunctive therapies, thereby personalizing and optimizing mental health care pathways.

发表机构

  • Deakin University(迪肯大学)
  • Black Dog Institute(黑狗研究所)
  • University of New South Wales(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

↑