arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越标记:临床框架弥合自杀风险测量中的审核差距

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Shreyas Krishnan, Gun Ahn, Jungjin Kim

arXiv 2609.06263首次发表:更新:

发表机构

University of California, Berkeley; Wondi AI; MIT; Harvard Medical School; McLean Hospital(加州大学伯克利分校; Wondi AI; 麻省理工学院; 哈佛医学院; 麦克莱恩医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出用临床分级框架替代二元标记来测量自杀风险,通过专家撰写的提示词显著提升审核API的严重程度评估准确性,并发布516条帖子基准数据集。

AI 中文摘要

审核API的设计初衷是标记违反政策的内容,而非衡量分级临床风险。但平台的义务不止于检测:对被动痛苦的回应与对有手段的主动计划的回应截然不同,而新兴法规(如加州参议院第243号法案)正将这一区分转化为合规要求。因此,我们探究已部署的安全信号能在多大程度上恢复具有临床意义的严重程度。我们发布了一个包含516条r/SuicideWatch帖子的基准数据集,由持照精神科医生依据哥伦比亚自杀严重程度评定量表,按四级有序量表(指标、意念、行为、尝试)进行评级,并在七项有序感知指标下评估审核API、提示大语言模型和监督基线。三项发现:供应商审核API能很好地区分低严重程度与高严重程度帖子(高风险F1为0.860),但对严重程度的测量效果较差(宏F1为0.395),系统性地过度预测最严重类别。基于临床的零样本提示恢复了大部分差距(宏F1为0.562),而专家撰写的框架(而非微调、增加推理或朴素多智能体聚合)是有效的杠杆。推理的价值取决于语境:在冗长、嘈杂的Reddit帖子上有损性能,在简短、临床医生撰写的陈述上有助性能。我们认为,分级严重程度而非二元标记,才是相称注意义务所要求的,并发布我们的评估框架以支持该测量。

英文摘要

Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emerging regulation (e.g., California Senate Bill 243) is turning that distinction into a compliance requirement. We therefore ask how well deployed safety signals recover clinically meaningful severity. We release a benchmark of 516 r/SuicideWatch posts rated by a licensed psychiatrist on a four-level ordinal schema (Indicator, Ideation, Behavior, Attempt) grounded in the Columbia Suicide Severity Rating Scale, and evaluate moderation APIs, prompted LLMs, and supervised baselines under seven ordinal-aware metrics. Three findings. Vendor moderation APIs separate low- from high-severity posts well (0.860 high-risk F1) but measure severity poorly (0.395 macro F1), systematically over-predicting the most severe category. Clinically grounded zero-shot prompting recovers much of that gap (0.562 macro F1), and expert-authored framing (not fine-tuning, added reasoning, or naive multi-agent aggregation) is the effective lever. The value of reasoning depends on register: it hurts on long, noisy Reddit posts and helps on short, clinician-authored statements. We argue graded severity, not a binary flag, is what a proportionate duty of care requires, and release our evaluation framework to support that measurement.

Comments20 pages, 6 figures. Accepted at the 5th Workshop on NLP for Positive Impact (NLP4PI), EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑