arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07766cs.CLcs.AI

技巧的宝库还是神话的宝库?在可解释的自杀风险评估中利用任务知识降低建模复杂度

Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

  • Indian AI Research Organisation(印度人工智能研究组织)
  • Ahmedabad University(艾哈迈达巴德大学)
  • University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

Shlok Shelat, Shrey Salvi, Souvik Roy, Manas Gaur, Amit Sheth

AI总结:

本研究通过大规模受控实验审计了常见NLP技巧在自杀风险评估中的有效性,发现仅少数可靠,并据此构建了任务导向系统,在三个输出上取得优异性能,提出任务条件技术选择原则。

AI中文摘要:

从社交媒体文本评估自杀风险是一个小数据、高风险的应用场景,不仅需要预测严重程度,还需要提供支持性证据以及临床相关的风险和保护因素。然而,常见的自然语言处理技术,包括模型扩展、合成数据、损失重加权、集成学习和阈值调整,往往在没有测试其收益在严重类别不平衡、耦合输出和有限的作者级数据下是否成立的情况下就被应用。我们研究了1,635条临床医生标注的帖子,并通过大约300个在作者不相交分区上的受控实验,审计了来自7个方法家族的31种预定义技术。我们发现,在此场景下,没有先前对该策略集的审计。研究结果指导了一个面向任务的系统,该系统输出三项结果:4级自杀风险、证据片段和24个临床风险及保护因素。31项比较中仅有5项产生了可靠的收益。我们将因素预测重新表述为每个帖子与其代码簿定义之间的蕴含关系,使用一个架构多样的集成,并进行类别平衡训练和分数重新缩放。风险预测条件化一个7模型证据标注器集成;证据限制符号风险规则;一个困难风险类别被单独路由。因素预测器保持独立,因为风险证据不提供额外的因素信号。我们还通过部署一致的校准,纠正了用于阈值拟合的验证分数与测试时集成分数之间的不匹配,这为因素系统带来了最大的改进。最终系统在风险上达到0.8203,在证据上达到0.7953,在因素上达到0.7045宏F1,综合得分为0.7781,在53个团队中排名第三。我们称其基本原则为任务条件技术选择:仅当任务特定知识、结构或经验证据证明其合理性时,才保留技术。

英文摘要:

Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidence and clinically relevant risk and protective factors. Yet common NLP techniques, including model scaling, synthetic data, loss reweighting, ensembling, and threshold tuning, are often applied without testing whether their gains hold up under severe class imbalance, coupled outputs, and limited author-level data. We study 1,635 clinician-annotated posts and audit 31 pre-specified techniques from 7 methodological families through roughly 300 controlled experiments on author-disjoint partitions. We found no prior audit of this playbook in this regime. The findings guide a task-grounded system for three outputs: 4-level suicide risk, evidence spans, and 24 clinical risk and protective factors. Only 5 of 31 comparisons produced reliable gains. We reformulate factor prediction as entailment between each post and its codebook definitions, using an architecturally diverse ensemble with class-balanced training and score rescaling. Risk predictions condition a 7-model evidence tagger ensemble; evidence restricts symbolic risk rules; and a difficult risk class is routed separately. The factor predictor remains independent because risk evidence provides no additional factor signal. We also correct a mismatch between validation scores used for threshold fitting and test-time ensemble scores through deployment-consistent calibration, yielding the largest improvement to the factor system. The final system achieves 0.8203 for risk, 0.7953 for evidence, and 0.7045 macro-F1 for factors, with a 0.7781 composite, ranking third among 53 teams. We call the underlying principle task-conditioned technique selection: retain techniques only when task-specific knowledge, structure, or empirical evidence justifies them.

补充信息

↑