发表机构
Keido Labs(启道实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究发现LLM作为评判器评估对话式AI心理安全性存在缺陷,提出经心理学家校正的本地模型aipsy-judge-1.0,其危机检测等性能优于前沿模型,且评分更忠实、数据可本地留存。
AI 中文摘要
将前沿大语言模型(LLM)用作评判器的标准方法——挑选一个前沿模型或对多个模型取平均——在评估对话式AI的心理安全性时存在明显的不安全问题。我们使用开源的冻结安全工具aipsy-bench开展了一项完全交叉的能力研究:选取三个前沿模型(gpt-5.4-mini、claude-sonnet-4-6、gemini-2.5-flash)作为生成器和评判器,针对3000条心理健康、陪伴及辅导类消息,与心理学家的评分进行对比。评判分歧并非随机噪声,而是具有结构性,集中在安全关键指标上,其中Gemini是异常值——它最为宽松,带有+0.99的自我偏好溢价,标记的极端失败案例少得多,还将一条涉及“手边有工具的自我伤害”的响应评为“ exemplary( exemplary 此处指符合要求的范例)”。在存在谄媚行为的同理心维度上,评判者间的一致性最低(α=0.24);而二元危机检测标记是评判者唯一达成一致的安全关键信号(α=0.80),倾向于过度标记,这对于分诊筛查而言是安全方向。作为标准解决方案的等权重平均,会将这种宽松性和对极端失败的忽视融入安全评分。现成的开放权重评判器表现更差,原因在于其倾向性而非能力,而倾向性是可微调的。因此,我们将针对每个指标的、经心理学家校正的目标蒸馏为一个小型冻结本地模型aipsy-judge-1.0,它是基于Gemma-4-26B-A4B的Apache-2.0许可微调模型。aipsy-judge-1.0在综合指标(组内相关系数ICC从0.64提升至0.75)和危机检测(Cohen’s kappa从0.65提升至0.82)上比其基础模型更贴合校正后的目标,能检测92%的危机且倾向于少报假阳性,比任何单个前沿评判器的评分都更忠实,同时所有文本都保留在本地设备上。这些是针对单一专家指导目标的定向读数,未经过多评判者间一致性验证,与供应商后训练模型共享安全评判器会存在相同的盲区。
英文摘要
The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response "exemplary." Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree on (alpha 0.80), erring toward over-flagging, the safe direction for a triage screen. Equal-weight averaging, the canonical fix, blends that leniency and tail-blindness into the safety score. Off-the-shelf open-weight judges are worse for a dispositional, not capability, reason -- and disposition is fine-tunable. We therefore distill a per-metric, psychologist-corrected target into a small, frozen, local model, aipsy-judge-1.0, an Apache-2.0 fine-tune of Gemma-4-26B-A4B. aipsy-judge-1.0 tracks the corrected target better than its base on the composite (ICC 0.64 to 0.75) and crisis detection (kappa 0.65 to 0.82), catches 92% of crises with a false-positive lean, and grades more faithfully than any single frontier judge, while every transcript stays on the machine. These are directional readings against a single-expert-informed target, not validated multi-rater agreement. A safety grader that shares a vendor's post-training shares its blind spots.
Comments24 pages, 3 figures, 6 tables. Model available at https://huggingface.co/keidolabs/aipsy-judge-1.0 (Apache-2.0)