发表机构
Singapore University of Technology and Design(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对AI生成教育插件的合规性问题,提出EduPluginBench基准与分阶段准入方法,经实验验证其可显著提升缺陷召回率,但编码模型生成结果的迁移符合性表现不佳。
AI 中文摘要
代码生成模型可产出可执行组件,但编译与功能测试无法验证其是否符合最小权限、遥测同意、溯源、特权写入权限、生命周期约束或有限失效等要求。本文提出EduPluginBench,这是面向受管控软件生态中生成插件的可执行基准与分阶段准入方法。针对30个规格的1440个经激活检查的一阶变异体,P0-P4相较于P0-P2,发布阻塞缺陷召回率提升74.7个百分点(规格聚类95%置信区间为73.4-75.8),且120个干净参考中未观测到拒绝(95%威尔逊上界为3.1%)。对两个当前编码模型的600个未修改生成结果开展冻结迁移研究,发现300/600个可解析,但无一个通过P0或达到P0-P4符合性(95%上界为0.64%),下游保障估计值未定义。一项独立标注的Moodle研究保留了16个易受攻击/修复对,冻结通用PHP检测器未发现易受攻击修订。这些负迁移结果避免将受控契约一致性误读为独立真实缺陷有效性。此前540个生成结果的诊断显示,事后有限修复产生112个P0通过,全部不符合要求,召回率从13.4%提升至100%。该工件保留协议、公开源码溯源、原始生成结果、行级决策、审计、分析代码及复现说明。
英文摘要
Code-generation models can produce executable components, but compilation and functional tests do not establish compliance with least privilege, telemetry consent, provenance, privileged-write authority, lifecycle constraints, or bounded failure. We introduce EduPluginBench, an executable benchmark and staged admission method for generated plugins in governed software ecosystems. Across 1,440 activation-checked first-order mutants from 30 specifications, P0-P4 increased release-blocking-defect recall by 74.7 percentage points (specification-clustered 95% CI 73.4-75.8) over P0-P2, with no observed rejection among 120 clean references (95% Wilson upper bound 3.1%). A frozen transfer study of 600 unmodified generations from two current coding models found that 300/600 parsed, but none passed P0 or achieved P0-P4 conformance (95% upper bound 0.64%); downstream assurance estimands were undefined. An independently labelled Moodle study retained 16 vulnerable/fixed pairs; the frozen generic PHP detector found no vulnerable revisions. These negative transfer results prevent controlled contract consistency from being read as independent real-defect effectiveness. An earlier 540-generation diagnostic found that post-hoc bounded repair yielded 112 P0 passes, all nonconforming, with recall increasing from 13.4% to 100%. The artifact retains protocols, public-source provenance, raw generations, row-level decisions, audits, analysis code, and reproduction instructions.
Comments27 pages, 3 figures, 8 tables. Submitted to ACM Transactions on Software Engineering and Methodology (TOSEM)