arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20353cs.CLcs.AI

分歧假说:揭示心理健康自然语言处理中的词汇干扰与标签偏差

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

  • Doha Institute for Graduate Studies(多哈高等研究院)

机构由 AI 辅助整理,请以论文原文为准。

Moustafa Yehia Hassan

AI总结:

该研究提出TSS诊断框架与DoD统计量,揭示心理健康NLP分类器的词汇干扰与标签偏差,可标记特定标签源的捷径学习,为相关模型审计提供支撑。

AI中文摘要:

计算心理健康(CMH)分类器在分布偏移下常性能下降,原因是人工标注者与远程监督流程会奖励不同的语言信号。我们提出TSS(Triple-Stream Stress probe,三流压力探测仪),这是一种多通道诊断框架,将文本分解为:(A)词汇字符n元组、(B)小型、主要为无内容的形态句法通道、(C)含154个特征的心理语言学风格通道。在四个英文数据集(N=12906)上,TSS揭示了词汇干扰效应:向风格通道添加词汇特征会降低人工标注数据的Macro-F1(平均下降0.072,p<10^-4),但对自动标注数据无此影响。我们提出分歧度(DoD),一种改编自计量经济学的双重差分统计量,用于标签源审计,兼具实例级自助法推断;核心估计值为DoD(BC-A)=0.0374,95%置信区间[0.0097,0.0651],p=0.0032。仅针对Twitter的平台分层DoD(消除Reddit与Twitter的对比)通过自助法推断重现该模式:DoD-Tw(BC-A)=+0.096(p<0.001),DoD-Tw(AC-A)=-0.089(p<0.001)。干预性掩码(pos_only)在破坏人工数据集内容词后,保留了约95-99%的C通道性能,表明风格通道不主要依赖词汇表层形式。TSS被定位为诊断审计框架,而非临床筛查工具:它在提出泛化主张前,标记特定标签源的捷径学习。

英文摘要:

Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a multi-channel diagnostic framework that decomposes text into (A) lexical character n-grams, (B) a small, mostly content-free morpho-syntactic channel, and (C) a 154-feature psycholinguistic style channel. Across four English datasets (N=12,906), TSS reveals a lexical interference effect: adding lexical features to the style channel reduces Macro-F1 on human-labeled data (mean drop 0.072, p<10^-4) but not on auto-labeled data. We propose Degree of Divergence (DoD), a difference-in-differences statistic adapted from econometrics for label-source auditing, with instance-level bootstrap inference; the headline estimate is DoD(BC-A) = 0.0374, 95% CI [0.0097, 0.0651], p=0.0032. A platform-stratified Twitter-only DoD (which removes the Reddit vs. Twitter contrast) reproduces the pattern with bootstrap inference: DoD-Tw(BC-A) = +0.096 (p<0.001) and DoD-Tw(AC-A) = -0.089 (p<0.001). Interventional masking (pos_only) retains ~95-99% of Channel C's performance after destroying content words on human datasets, indicating that the style channel does not rely primarily on lexical surface form. TSS is positioned as a diagnostic audit framework, not a clinical screening tool: it flags label-source-specific shortcut learning before generalization claims are made.

补充信息

↑