发表机构
Weill Cornell Medicine; Cornell University(威尔康奈尔医学院; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于QLoRA的多任务系统,用于社交媒体可解释自杀风险评估,通过联合训练与定制聚合策略,在IEEE BigData 2026 Cup上取得综合得分0.7738。
AI 中文摘要
可解释的自杀风险评估要求模型不仅估计风险严重程度,还要识别支持性语言以及帖子中表达的风险和保护因素。我们介绍了为IEEE BigData 2026 Cup社交媒体可解释自杀风险评估竞赛开发的系统,该系统处理三个任务:风险等级分类、证据短语提取和多标签因素识别。我们的方法采用量化低秩适配(QLoRA)和答案掩码的因果语言模型目标来适配Qwen2.5-Instruct模型。我们针对风险分类联合训练所有三个任务,针对证据提取联合训练任务1a和1b,并单独适配任务2进行因素识别。我们还针对每个输出定制聚合策略:我们平均32B和72B模型的风险等级概率,通过交叉折叠共识合并证据短语,并基于折叠外操作点通过比率匹配校准因素特定决策。在官方排行榜上,最终系统综合得分为0.7738,任务1得分为0.8089,任务2得分为0.6919。在评估的配置中,三任务训练在任务1a上表现最佳,任务1a和1b的联合训练在任务1b上表现最佳,而任务特定训练在任务2上表现最佳。当组件模型具有互补错误时,概率平均进一步改善了任务1a。这些发现强调了在统一语言模型框架内,根据每个任务的输出结构定制训练目标和聚合策略的价值。
英文摘要
Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three tasks: risk-level classification, evidence phrase extraction, and multi-label factor identification. Our approach adapts Qwen2.5-Instruct models using quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. We jointly train across all three tasks for risk classification, jointly train on Tasks 1a and 1b for evidence extraction, and adapt Task 2 separately for factor identification. We also tailor aggregation to each output: we average risk-level probabilities from the 32B and 72B models, combine evidence phrases through cross-fold consensus, and calibrate factor-specific decisions through rate matching based on out-of-fold operating points. On the official leaderboard, the final system achieved a composite score of 0.7738, with 0.8089 on Task 1 and 0.6919 on Task 2. Across the evaluated configurations, three-task training performed best for Task 1a, joint training on Tasks 1a and 1b performed best for Task 1b, and task-specific training performed best for Task 2. Probability averaging further improved Task 1a when component models had complementary errors. These findings highlight the value of tailoring both training objectives and aggregation strategies to the output structure of each task within a unified language-model framework.