发表机构
University of the Cumberlands(坎伯兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在评估前沿人工智能代理能否自主进行临床人工智能安全审计。基于METR任务标准构建评估任务,让代理按要求实施攻击、计算得分并编写报告。通过对三个前沿模型在不同数据集和模型架构上的评估,得出各模型表现及成本等情况,部分成果已公开。
AI 中文摘要
临床人工智能模型若未检测到对抗性漏洞,可能会对患者造成伤害,但正式的安全审计需要统计专业知识、专门工具和大量时间。我们提出了一项基于METR任务标准v0.3.0的开放评估任务,测试前沿人工智能代理能否自主实施结构化临床人工智能安全审计。给定预训练临床预测模型、患者数据集和书面说明,各代理须从伪代码实施四种攻击,计算涵盖FGSM鲁棒性、成员推理抗性、预期校准误差和边界攻击抗性的安全态势得分,并在Docker容器中编写结构化JSON报告。六个变体跨越威斯康星州诊断乳腺癌和MIMIC-IV重症监护病房死亡率数据集,涉及三种防御强度递增的模型架构,参考分数从55.60到90.41。我们对三个前沿模型进行了54次评估,每个变体运行三次。Claude Sonnet 4.6和GPT-4.1完成了所有18次运行并获得完美评估分数。GPT-4o完成了61%的运行,每次运行使用的令牌数约为Claude的五倍。GPT-4.1的总API成本为8美元,Claude Sonnet 4.6为12美元,GPT-4o为27美元。GPT-4o的失败包括会话过早终止、聚合错误和提交文件为空。该任务、评分基础设施和威斯康星乳腺癌资产已公开发布;MIMIC-IV变体需要单独访问PhysioNet。
英文摘要
Clinical AI models can expose patients to harm when adversarial vulnerabilities go undetected, yet formal security auditing requires statistical expertise, specialized tools, and significant time. We present an open evaluation task, built on METR Task Standard v0.3.0, that tests whether frontier AI agents can autonomously implement a structured clinical AI security audit. Given a pre-trained clinical prediction model, a patient dataset, and written instructions, each agent must implement four attacks from pseudocode, compute a Security Posture Score covering FGSM robustness, membership inference resistance, expected calibration error, and boundary attack resistance, and write a structured JSON report in a Docker container using only a bash interface and no scaffolding code. Six variants span the Wisconsin Diagnostic Breast Cancer and MIMIC-IV ICU mortality datasets across three model architectures with increasing defense strength, with reference scores from 55.60 to 90.41. We ran 54 evaluations across three frontier models, with three runs per variant. Claude Sonnet 4.6 and GPT-4.1 completed all 18 runs and received perfect evaluator scores. GPT-4o completed 61 percent of runs and used about five times the per-run token count of Claude, although provider tokenization differs. Total API costs were 8 US dollars for GPT-4.1, 12 US dollars for Claude Sonnet 4.6, and 27 US dollars for GPT-4o. GPT-4o failures involved premature session termination, an aggregation error, and an empty submission file. The task, scoring infrastructure, and Wisconsin Breast Cancer assets are publicly released; MIMIC-IV variants require separate PhysioNet access.
Comments29 pages, 2 figures, 7 tables