arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22767cs.CLcs.LG

超越最终令牌分类:用于证据支持的自杀风险检测的异构读出

Beyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk Detection

Zirui Li, Yanling Li, Kaolanglang Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出异构读出分解(HRD)方法,通过分离语义验证与输出实现,利用Qwen3.8-27B模型和QLoRA适配器,在自杀风险检测基准上提升风险分类、因素检测和证据提取的性能,并识别出答案令牌边界是错误的关键来源。

中文摘要 AI 辅助

IEEE BigData Cup基准测试结合了三个具有不同输出结构的预测问题:有序自杀风险分类、多标签心理社会因素检测以及支持短语的提取。我们引入了异构读出分解(HRD),它将语义验证与输出实现分离。一个本地部署的Qwen3.8-27B模型,通过任务特定的QLoRA适配器进行调整,为基于卡片的查询生成答案令牌边际和第63层答案状态。HRD对有序风险比较四个潜在分数,保留大多数因素的令牌边际,同时将七个标签通过一个共享的潜在探针进行路由,并从逐字跨度候选中构建带有校准的、风险条件约束的证据集。在两个保留的用户分组确认折叠上,潜在风险读出将加权F1从0.8237提高到0.8372,宏F1从0.7965提高到0.8185。选择性因素路由将宏F1提高了0.0105,尾部标签宏F1提高了0.0189;相比之下,全局潜在替换和独立的标签特定探针失败。受约束的证据解码器在行级折叠外评估中将合并短语F1从0.7488提高到0.7609。Lenormand的最佳公开结果在子任务1上为0.8052,在子任务2上为0.6636,综合得分为0.7627。这些结果确定了答案令牌边界,而非仅语义表示,是该基准中可测量的错误来源。

英文摘要

The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label psychosocial factor detection, and extraction of supporting phrases. We introduce heterogeneous readout decomposition (HRD), which separates semantic verification from output realization. A locally deployed Qwen3.8-27B model, adapted with task-specific QLoRA adapters, produces both answer-token margins and layer-63 answer states for card-conditioned queries. HRD compares four latent scores for ordinal risk, retains token margins for most factors while routing seven labels through one shared latent probe, and constructs evidence sets from verbatim span candidates with calibrated, risk-conditional constraints. On two held-out user-grouped confirmation folds, the latent risk readout improves weighted F1 from 0.8237 to 0.8372 and macro F1 from 0.7965 to 0.8185. Selective factor routing improves macro F1 by 0.0105 and tail-label macro F1 by 0.0189; in contrast, global latent replacement and independent label-specific probes fail. The constrained evidence decoder raises pooled phrase F1 from 0.7488 to 0.7609 in row-level out-of-fold evaluation. Lenormand's best public result is 0.8052 on Subtask 1 and 0.6636 on Subtask 2, giving a composite score of 0.7627. These results identify the answer-token boundary, rather than semantic representation alone, as a measurable source of error in this benchmark.

发表机构

  • The Hong Kong Polytechnic University(香港理工大学)
  • Shenzhen University(深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑