心理健康自然语言处理中的跨平台泛化失效:对社交媒体上Transformer模型的五维度公平性审计
Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media
浏览论文内容
中文总结 AI 辅助
本研究提出CPFE五维度公平性审计框架,发现心理健康NLP的Transformer模型跨平台泛化失效显著,温度缩放可缓解校准问题,需将CPFE验证作为心理健康NLP系统的标准要求。
中文摘要 AI 辅助
我们提出了跨平台公平性评估(CPFE)框架——这是一项涵盖判别性能、校准度、统计显著性、预测公平性和归因稳定性的五维度审计协议,并将其应用于在Kaggle心理健康语料库(样本量n=35556)上训练的四个Transformer模型(BERT、RoBERTa、Emotion-DistilRoBERTa、GoEmotions-RoBERTa),并在Reddit(样本量n=6257)和Twitter(样本量n=2883)测试集上进行评估,其中情绪标签被映射为临床代理指标。所有三个独立评估的模型在跨平台上均表现出一致且显著的AUC下降:相对于平台内性能(AUC为0.983-0.987),Reddit上的AUC下降幅度为30.3%-35.4%,Twitter上为37.9%-39.5%,这一结果在五个独立训练种子中均得到验证。校准失效同时发生且极为严重:预期校准误差(ECE)从域内的0.056-0.060上升至Reddit的0.196-0.229和Twitter的0.499-0.542。平台特定的温度缩放可将平均ECE降低88.0%,且不改变判别性能(平均|ΔAUC|<0.01),证实存在可分离的失效模式。预测公平性分析显示跨平台存在巨大差异:原始差异指数(DI)<0.17;经先验偏移调整后的DI在Reddit上为0.11-0.29,Reddit中心理健康代理类别的均等机会差异为0.753-0.830,Twitter中焦虑类别的均等机会差异为0.755-0.831。归因稳定性分析显示跨平台存在近乎完全的词汇差异(在K=10时,16个模型-类别对中有14个的Jaccard相似度J=0)。这些发现支持将所有五个CPFE维度的跨平台验证作为异构环境中心理健康NLP系统的标准要求。在单种子微调实验中,平均AUC提升了0.216,表明目标平台标签作为训练信号比作为校准信号能提供更大的收益。
英文摘要
We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and apply it to four transformer models (BERT, RoBERTa, Emotion-DistilRoBERTa, GoEmotions-RoBERTa) trained on a Kaggle mental health corpus (n=35,556) and evaluated on Reddit (n=6,257) and Twitter (n=2,883) test sets with emotion labels mapped to clinical proxies. All three independently evaluated models exhibit consistent and substantial cross-platform AUC degradation (30.3-35.4% on Reddit, 37.9-39.5% on Twitter) relative to within-platform performance (AUC 0.983-0.987), confirmed across five independent training seeds. Calibration failure is concurrent and severe: ECE rises from 0.056-0.060 in-domain to 0.196-0.229 on Reddit and 0.499-0.542 on Twitter. Platform-specific temperature scaling reduces mean ECE by 88.0% without altering discriminative performance (mean |delta AUC|<0.01), confirming separable failure modes. Prediction equity analysis reveals large cross-platform disparities (raw DI < 0.17; prior-shift-adjusted DI: 0.11-0.29 on Reddit), with equalized odds differences of 0.753-0.830 for mental health proxy classes on Reddit and 0.755-0.831 for anxiety on Twitter. Attribution stability analysis shows near-complete vocabulary divergence across platforms (Jaccard J=0 in 14/16 model-class pairs at K=10). These findings support treating cross-platform validation across all five CPFE axes as a standard requirement for mental health NLP systems in heterogeneous environments. In a single-seed fine-tuning experiment, mean AUC improved by 0.216, suggesting target-platform labels provide greater benefit as training signal than as calibration signal.
发表机构
- Gyan Ganga Institute of Technology and Sciences(gyan ganga 技术与科学学院)
机构由 AI 辅助整理,请以论文原文为准。