AI 中文总结
本文提出用于围产期心理健康支持的安全门控多模态AI后端Anian,其通过分层状态表示与保守风险融合实现安全门控,原型在多任务评估中表现优异,验证了框架内部可行性。
AI 中文摘要
安全关键的心理健康支持系统必须区分何时适合进行支持性对话,何时应阻止自由形式的生成。本文提出Anian,这是一个用于围产期心理健康支持和正念干预路由的安全门控多模态AI后端。Anian不旨在诊断精神疾病,也不替代临床护理或危机干预。其模块化流程将生成式AI置于结构化状态表示、保守风险融合和响应门控的下游。用户文本或语音衍生的ASR转录本被映射到四个关联层级:L1情绪状态、L2社会心理构念、L3安全风险和L4干预路由。使用最高风险优先级规则融合本地文本和基于规则的安全证据与外部语音衍生证据,融合风险公式为S_fusion = max(S_local, S_external)。在中等或高融合风险下,普通AI生成的响应和文本转语音交付被阻止,取而代之的是固定安全内容和人工支持提示。内部原型评估使用了来自公开情绪、对话、心理健康相关及中文对话语料库的约858295条归一化记录,在弱标签和规则衍生框架内进行。L1情绪分类的微F1分数为0.9604,L2社会心理构念为0.9144,L4路由为0.9742。在对233个样本的受控安全压力测试中,L3规则引擎在预定义场景内实现了1.0000的高风险召回。这些发现支持标签框架和门控逻辑的内部可行性,但未确立临床有效性、诊断准确性、现实世界安全性或有效性。我们报告了架构、本体、安全融合机制、原型评估、错误分析计划以及专家审核和现实世界验证的路线图。
英文摘要
Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation should be blocked. This paper presents Anian, a safety-gated multimodal AI backend for perinatal mental-health support and mindfulness-intervention routing. Anian is not intended to diagnose psychiatric conditions or replace clinical care or crisis intervention. Its modular pipeline places generative AI downstream of structured state representation, conservative risk fusion, and response gating. User text or voice-derived ASR transcripts are mapped into four linked layers: L1 emotion states, L2 psychosocial constructs, L3 safety risk, and L4 intervention routes. Local text- and rule-based safety evidence is fused with external voice-derived evidence using a highest-risk-priority rule, S_fusion = max(S_local, S_external). At moderate or high fused risk, ordinary AI-generated responses and text-to-speech delivery are blocked and replaced by fixed safety content and prompts for human support. An internal prototype evaluation used approximately 858,295 normalized records from public emotion, dialogue, mental-health-related, and Chinese dialogue corpora within a weak-label and rule-derived framework. Micro-F1 scores were 0.9604 for L1 emotion classification, 0.9144 for L2 psychosocial constructs, and 0.9742 for L4 routing. In a controlled safety stress test of 233 samples, the L3 rule engine achieved high-risk recall of 1.0000 within predefined scenarios. These findings support the internal feasibility of the label framework and gating logic but do not establish clinical validity, diagnostic accuracy, real-world safety, or effectiveness. We report the architecture, ontology, safety-fusion mechanism, prototype evaluation, error-analysis plan, and roadmap for expert-reviewed and real-world validation.
Comments16 pages, 4 figures