arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09740cs.SE

安全测试作为LLM代码生成的可执行规范:益处、权衡与覆盖范围限制

Security Tests as Executable Specifications for LLM Code Generation: Benefits, Trade-offs, and Coverage Limits

Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出将安全测试作为LLM代码生成的可执行规范,开发SecTDD框架,通过多维度评估发现预先展示可见测试可提升联合成功率,结构化反馈修复效果更优但益处受多因素影响。

中文摘要 AI 辅助

大型语言模型(LLM)可生成功能有用但仍存在漏洞的代码,而以安全为重点的干预措施可能会破坏预期行为。我们研究将安全测试作为可执行规范,分别在代码生成前和迭代修复过程中应用。我们开发了SecTDD,这是一种受控的测试反馈框架,可分离三个因素:是否预先展示测试、失败的执行是否触发修订,以及如何选择和表示失败。评估采用行为划分的可见测试和隐藏测试,以及字节完全相同的初始候选代码用于修复比较。在2705条轨迹、31个任务实例、三个安全代码基准、16个CWE类别和两个模型家族的范围内,预先展示所有可见测试可使隐藏的功能与安全联合成功率平均提高19.3个百分点,但仅改善了9个基准-模型条件中的7个,且对2个条件产生了损害。在共享候选代码比较中,结构化反馈修复了80个初始不成功的候选代码,无联合退化;固定原始反馈修复了83个,但导致3个退化。结构化和原始反馈在其他方面几乎无法区分(6次获胜、6次失败和453次平局)。在所有常见机制下,通过所有可见测试的候选代码仍会在隐藏的行为族下失败。这些结果表明,可执行反馈可修复安全代码生成,但其益处取决于模型、任务、反馈切入点,尤其是测试覆盖范围。

英文摘要

Large language models (LLMs) can generate functionally useful code that remains vulnerable, while security-focused interventions may break intended behavior. We investigate security tests as executable specifications both before generation and during iterative repair. We develop SecTDD, a controlled test-feedback scaffold that separates three factors: whether tests are shown upfront, whether failed executions trigger revision, and how failures are selected and represented. The evaluation uses behavior-partitioned visible and hidden tests and byte-identical initial candidates for repair comparisons. Across 2,705 trajectories, 31 task instances, three secure-code benchmarks, 16 CWE categories, and two model families, showing all visible tests upfront increases hidden functional-and-security joint success by 19.3 percentage points on average, but improves only seven of nine benchmark-model conditions and harms two. In shared-candidate comparisons, structured feedback repairs 80 initially unsuccessful candidates with no joint regressions; fixed raw feedback repairs 83 but causes three regressions. Structured and raw feedback are otherwise nearly indistinguishable head-to-head (six wins, six losses, and 453 ties). Candidates that pass all visible tests still fail hidden behavior families under every common regime. These results show that executable feedback can repair secure-code generation, but its benefits depend on the model, task, feedback entry point, and especially test coverage.

↑