发表机构
Monash University; Orygen; The University of Melbourne; Swinburne University of Technology; University of Liverpool; Khalifa University(莫纳什大学; 奥瑞金; 墨尔本大学; 斯威本科技大学; 利物浦大学; 哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AnchorSIPS合成数据集,基于Mini-SIPS访谈构建含1万次结构化精神病风险访谈,解决临床数据隐私问题,用于研究证据提取等,实验发现LLM基线存在细节提取不足问题。
AI 中文摘要
精神病风险评估领域的AI进展受限于数据访问瓶颈,真实临床访谈因隐私、治理和知情同意限制难以共享。本文提出AnchorSIPS,这是一个包含10000次结构化精神病风险访谈的合成数据集,具有基于 transcript(转录文本)的测量目标。每次访谈以临床医生管理的精神病风险访谈Mini-SIPS为模型,涵盖病史、24个症状问题、患者确认项的跟进证据、类妄想症状(异常信念)、类幻觉症状(异常感知)和言语紊乱的决策、明确精神病水平症状(“显性精神病”)的排除,以及最终的轻度精神病综合征(APS)诊断——这是轻度或早期精神病症状的高风险状态。APS诊断不是独立标签,它依赖于早期确认、支持性跟进细节、症状类别决策和显性精神病检查,每个中间决策都锚定到其支持的转录文本轮次。AnchorSIPS由“规划-实现”流水线生成:隐藏的病例表指定患者临床状态,确定性规划器确定访谈结构,大语言模型(LLM)仅在验证和受限修复下生成患者话语。在生成前固定标签和结构,可避免多轮LLM对话中常见的轮间不一致。在7个LLM基线中,模型能恢复粗略决策,但无法提取跟进细节或引用支持性转录文本轮次,因此最终标签性能高估了访谈能力。AnchorSIPS旨在用于证据提取、基于转录文本的测量和部分披露下的不确定性研究。
英文摘要
Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.