arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

构建可信的Bluesky心理健康基准:一种验证感知的弱监督框架

Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework

Gaurab Chhetri, Anandi Dutta, Subasish Das

arXiv 2609.22696首次发表:更新:

发表机构

Texas State University(德克萨斯州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种验证感知的弱监督框架,在去中心化平台Bluesky上构建自杀意念与心理健康披露基准,整合数据采集、词典过滤、Llama-3-8B标注及人工验证,发现模型性能依赖任务定义与验证协议,并揭示弱标签的不同失败模式。

AI 中文摘要

去中心化社交媒体平台为计算心理健康研究带来了新的机遇和挑战,因为数据访问、内容审核、标注和部署责任分布在多个技术和治理层面。本文提出了一种验证感知的弱监督系统,用于在Bluesky(一个基于AT Protocol构建的去中心化社交媒体平台)上构建和评估自杀意念(SI)及更广泛的心理健康(MH)披露基准。该系统整合了公共firehose数据采集、任务特定的词典过滤、Llama-3-8B辅助的二分类标注、人工裁决的验证子集以及基于Transformer的模型基准测试。利用该流程,我们构建了两个任务特定的语料库,分别包含8,346条SI标注帖子和9,988条MH标注帖子。评估结果表明,模型性能强烈依赖于任务定义和验证协议。BERT+LSTM在SI分层交叉验证中取得了最高的F1分数,RoBERTa在SI留出集上取得了最强的F1分数,而DistilRoBERTa在MH交叉验证中取得了最佳的F1分数。人工验证揭示了不同任务中弱标签的不同失败模式,其中SI标签以假阴性为主,而MH标签以假阳性为主。这些发现表明,去中心化社交媒体可以支持可复现的心理健康基准测试,但前提是系统设计、标签来源、验证策略和部署约束必须同时进行评估。

英文摘要

Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities are distributed across multiple technical and governance layers. This paper presents a validation-aware weak-supervision system for constructing and evaluating suicidal ideation (SI) and broader mental health (MH) disclosure benchmarks on Bluesky, a decentralized social media platform built on the AT Protocol. The system integrates public firehose collection, task-specific lexicon filtering, Llama-3-8B-assisted binary annotation, human-adjudicated validation subsets, and transformer-based model benchmarking. Using this pipeline, we construct two task-specific corpora containing 8,346 SI-labeled posts and 9,988 MH-labeled posts. The evaluation shows that model performance depends strongly on both task definition and validation protocol. BERT+LSTM achieves the highest SI stratified cross-validation F1-score, RoBERTa achieves the strongest SI holdout F1-score, and DistilRoBERTa achieves the best MH cross-validation F1-score. Human validation reveals different weak-label failure modes across tasks, with SI labels dominated by false negatives and MH labels dominated by false positives. These findings show that decentralized social media can support reproducible mental health benchmarking, but only when system design, label provenance, validation strategy, and deployment constraints are evaluated together.

CommentsThis is the author's preprint version of a paper accepted for presentation at HICSS 60 (Hawaii International Conference on System Sciences), 2027, Hawaii, USA. The final published version will appear in the official conference proceedings. Conference site: https://hicss.hawaii.edu/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑