发表机构
Kinvectum AB(金维图姆公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在信息缺失下评估医疗人工智能,通过删除对话部分对四个模型压力测试,发现评判者选择影响表面安全性,语言模型评判者更宽松,安全差距在于校准,还发布了相关测试工具等。
AI 中文摘要
医疗人工智能的就绪压力测试主要集中在封闭式和多模态基准测试上。我们将其扩展到信息缺失情况下的开放式临床对话,其中安全行为意味着识别缺失信息并进行限定、澄清或不过度承诺,并且评估者成为测量的一部分。我们通过删除HealthBench对话中最终用户轮次的后半部分,对四个模型进行压力测试,这四个模型分别是三个旗舰模型(Claude Opus 4.8、GPT-5.5、Grok 4.3)和一个中级模型(Gemini 3.5 Flash),并由一个由四个提供者组成的语言模型评判小组和一个盲法临床医生锚定参考对回答进行评分。两个面向评估者的结果很稳健。首先,评判者的选择会显著改变表面安全性:评判者之间的一致性仅为中等(Fleiss' kappa = 0.65),在调整每个评判者的总体宽松度(投票级逻辑回归)后,同提供者的正向关联仍然存在(精确排列p = 0.04;GPT-5.5在概率尺度上约为 +0.10),大到足以改变一旦排除其自身提供者评判者后哪个模型似乎过度承诺最少。其次,在一个盲法的50项子样本上,语言模型评判者比临床医生更宽松:所有四个模型都比更严格的独立临床医生显著更宽松(在66 - 84%的项目上认可适当的不确定性,而临床医生为52%),四个模型中的三个比受作者影响的共识更宽松(仅Grok有方向性;评判者与共识的kappa = 0.20 - 0.43)。在作者审核的临床未确定子集中,宽松度差距扩大,点估计模型排序保持不变。一个封闭式的MedQA锚定证实,四个模型中的三个模型的准确性很高,选项顺序效应在+/-5分的等效区域内,所以安全差距大约在于校准而非知识。我们发布了测试工具、提示、每项输出、评判小组、扰动审核和人工注释协议。
英文摘要
Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss' kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.
CommentsGitHub https://github.com/KAVentures/health-ai-readiness-robustness