AI能否编写合规代码?以及在多大程度上能?评估Claude Fable 5、Claude Opus 4.8和Claude Opus 5在四个用例中的SOC 2合规性
Can AI Write Compliant Code, and to What Extent? Evaluating SOC 2 Compliance of Claude Fable 5, Claude Opus 4.8, and Claude Opus 5 Across Four Use Cases
AI总结:
该研究评估Claude Fable 5、Claude Opus 4.8和Claude Opus 5在四个用例中的SOC 2合规性,发现添加一句SOC 2语句可大幅提升合规率,模式匹配评分器不可靠需替换为语义检查。
AI中文摘要:
软件团队如今将生产代码委托给语言模型,包括配置存储、处理凭证和存储受监管数据的代码,因此我们研究:当未提及安全性时,模型是否应用SOC 2程序期望的控制措施(加密、受限访问、日志记录、保留),以及仅添加一句提及该标准的话会如何改变结果。我们测试了三个前沿模型(Claude Fable 5、Claude Opus 4.8和Claude Opus 5)在四个用例(S3 CLI、身份验证服务、RDS Terraform模块、存储个人数据的文件上传处理器)中的表现,每个用例分别从中性任务语句生成一次,以及添加一句包含SOC 2的语句后生成一次,共24个输出,我们对照映射到特定信任服务标准的二元评分标准对其进行评分,并手动验证所有失败和标记的行为。未提示时的合规率为47%至88%,这取决于控制措施是否是代码常规编写的一部分,因此密码哈希和storage_encrypted在未被要求时会出现,而S3加固、保留和MFA钩子则不会。中性提示还引入了实际漏洞,包括可访问的Werkzeug调试器允许远程代码执行、未认证下载、返回所有存储姓名和电子邮件的端点,这些都被我们的第一个评分标准评为合格,第四个缺陷通过是因为其值由条件计算得出。添加一句SOC 2语句后,所有用例的合规率提升至86%至100%,分值增加23至50分,并消除了所有不安全结构,不过模型任务概念之外的控制措施仍存在,包括MFA钩子、cookie标志和账户生命周期。模型选择的影响最小,同代模型在所有八个单元的评分标准项中差异不超过一项,而模式匹配评分器不可靠,在216次判断中有27次与语义分级不一致,且放过了一个实际缺陷,因此需要替换为语义检查。
英文摘要:
Software teams now delegate production code to language models, including code that provisions storage, handles credentials, and stores regulated data, so we asked whether a model applies the controls a SOC~2 program expects (encryption, restricted access, logging, retention) when nobody mentions security, and how much one sentence naming the standard changes the answer. We tested three frontier models (Claude Fable~5, Opus~4.8, and Opus~5) across four use cases (an S3 CLI, an authentication service, an RDS Terraform module, and a file-upload handler holding personal data), each generated once from a neutral task statement and once with a single SOC~2 sentence added, scoring all 24 outputs against binary rubrics mapped to specific Trust Services Criteria and hand-verifying every failure and flagged act. Unprompted conformance ran from 47\% to 88\% and tracked whether a control is part of how the code is normally written, so password hashing and \texttt{storage\_encrypted} appear unasked while S3 hardening calls, retention, and MFA hooks do not. The neutral prompt also shipped real vulnerabilities, including a reachable Werkzeug debugger allowing remote code execution, an unauthenticated download, and an endpoint returning every stored name and email, all scored clean by our first checklist, with a fourth defect passing because its value was computed by a conditional. One SOC~2 sentence moved every case to 86--100\%, worth 23 to 50 points, and removed every insecure construction, though controls outside the model's conception of the task survived it, including MFA hooks, cookie flags, and account lifecycle. Model choice mattered least, with same-generation models within one rubric item across all eight cells, and the pattern-matching scorer proved unreliable, disagreeing with semantic grading on 27 of 216 judgments and passing a real defect, so it needs replacing with semantic checks.