将法规转化为代码:一种用于增强软件工程中大型语言模型(LLM)选择的治理与合规性的模型
Operationalizing Regulations into Code: A Model to Enhance Governance and Compliance in LLM Selection for Software Engineering
AI总结:
本文提出一种基于设计科学研究的三层模型,将法规转化为可操作的LLM选择标准,试点评估显示该模型可识别合规风险,为软件工程LLM治理提供可行方案。
AI中文摘要:
将大型语言模型(LLMs)集成到软件开发生命周期(SDLC)中可提升开发者生产力,但也会在模型选择阶段引入安全、隐私和合规风险。欧盟人工智能法案(EU AI Act)、美国国家标准与技术研究院人工智能风险管理框架(NIST AI Risk Management Framework, RMF)、通用数据保护条例(GDPR)、巴西通用数据保护法(Lei Geral de Proteção de Dados, LGPD)以及ISO/IEC 42001等法规和框架设定了相关义务,这些义务往往难以转化为技术决策的可操作标准。本文提出一种用于支持软件工程项目中LLM选择的治理与合规性的模型,该模型通过设计科学研究(Design Science Research, DSR)开发,分为三个层次:(i)监管要求层;(ii)组织治理能力层,通过包含淘汰标准(knock-out criteria)和加权评分标准的多标准决策矩阵实例化;(iii)生产力与可持续性结果层,通过LLM治理评估协议(Protocol for LLM Governance, PAG-LLM)实现。一个监管反馈循环将操作结果反馈回规范层,使模型能够迭代优化。基于通用弱点枚举(Common Weakness Enumeration, CWE)和OWASP Top 10的20个对抗场景的试点评估表明,商业云LLMs与本地开源LLMs存在不同的风险特征。结果提供了初步证据,即监管淘汰逻辑(尤其是淘汰标准)可阻止选择技术上有竞争力但合规风险不可接受的模型,证明了面向治理的软件工程项目LLM选择的可行性。
英文摘要:
Integrating Large Language Models (LLMs) into the Software Development Life Cycle (SDLC) can improve developer productivity, but it also introduces security, privacy, and compliance risks during model selection. Regulations and frameworks such as the EU AI Act, the NIST AI Risk Management Framework (RMF), the General Data Protection Regulation (GDPR), the Lei Geral de Proteção de Dados (LGPD), and ISO/IEC 42001 establish obligations that are often difficult to translate into operational criteria for technical decision-making. This paper proposes a model to support governance and compliance in LLM selection for software engineering projects. The model is developed through Design Science Research (DSR) and is structured in three layers: (i) regulatory requirements, (ii) organizational governance capabilities, instantiated by a multi-criteria decision matrix with knock-out and weighted scoring criteria, and (iii) productivity and sustainability outcomes, operationalized by the LLM governance assessment protocol (PAG-LLM). A regulatory feedback loop connects operational results back to the normative layer, enabling iterative refinement of the model. A pilot evaluation with 20 adversarial scenarios based on Common Weakness Enumeration (CWE) and the OWASP Top 10 suggests distinct risk profiles between commercial cloud-based LLMs and local open-source LLMs. The results provide preliminary evidence that regulatory disqualification logic, particularly K.O. criteria, can prevent the selection of technically competitive models that nonetheless pose unacceptable compliance risks, demonstrating the feasibility of governance-oriented LLM selection in software engineering projects.