arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13742cs.SEcs.AIcs.LG

基于ISO标准的非功能性需求(NFR)规范是否能提升大语言模型(LLM)的代码生成能力?针对丰富干预方式、结构化干预方式与自然语言基线的对比研究

Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline

Joào Pedro Monteiro Pereira, Vinicius Cardoso Garcia

首次发表
浏览论文内容

中文总结 AI 辅助

该研究对比了ISO标准下丰富自然语言、结构化JSON与单行基线三种NFR规范对LLM代码生成的影响,发现前者可提升代码静态质量,且语义内容比格式更重要。

中文摘要 AI 辅助

在基于大语言模型(LLM)的代码生成任务中,非功能性需求(NFR)常被表述为简短的单行短语。本研究探究将这些需求规范基于ISO/IEC 25010质量模型进行标准化处理,分别采用丰富自然语言文本(NL-rich)或结构化JSON(Structured)形式,是否能比RobuNFR风格的单行基线(NL-simple)在HumanEval、HumanEval-ET数据集上生成更优代码。研究在固定模型快照下,对四种NFR(性能、异常处理、代码异味、可读性)各设置十个提示变体,采用配对非参数分析进行评估。主要发现:基于ISO标准的丰富化处理提升了静态质量代理指标(四种NFR下的不可读性密度均下降,例如性能指标从NL-rich的0.88降至0.69),并降低了对提示措辞的敏感性,但未可靠提升功能正确性;其中异常处理维度的扩展测试通过率下降,表明防御性编码模式与精确输出基准之间存在张力。次要发现:当ISO内容保持恒定时,NL-rich与Structured在正确性上的差异可忽略不计(|Δ|≤0.023),说明语义内容比JSON与文本的格式更重要。研究结论:从业者应投入资源构建基于标准的NFR内容,而非纠结于序列化形式。研究提供了完全可追溯的复现包。

英文摘要

In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line phrases. We ask whether grounding those specifications in ISO/IEC 25010 Quality Model, either as rich natural-language prose (NL-rich) or as structured JSON (Structured), improves code generated on HumanEval/HumanEval-ET compared to a RobuNFR-style one-line baseline (NL-simple). We evaluate four NFRs (performance, error handling, code smell, readability) with ten prompt variations per condition under a fixed model snapshot and paired non-parametric analysis. Primary finding: ISO-grounded enrichment improves static quality proxies (unreadability density falls across all four NFRs (e.g., Performance 0.88 -> 0.69 for NL-rich)) and reduces sensitivity to prompt wording, but does not reliably improve functional correctness; for error handling, extended-test pass rate decreases, suggesting tension between defensive coding patterns and exact-output benchmarks. Secondary finding: when ISO content is held constant, NL-rich and Structured differ negligibly in correctness (|delta| <= 0.023), indicating that semantic content matters more than JSON-vs-prose format. Practitioners should invest in standard-grounded NFR content rather than serialization form. A fully traceable replication package is provided.

发表机构

  • Centro de Informática, Universidade Federal de Pernambuco(伯南布哥联邦大学信息中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑