arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

两个真相还是一个谎言?对现成大语言模型进行需求质量评估的基准测试:性能、误报与漏报

Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses

Jannatul Shefa, Alejandro Salado, Paul Wach, Taylan G. Topcu

arXiv 2609.03230首次发表:更新:

发表机构

Virginia Tech; University of Arizona(弗吉尼亚理工大学; 亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过基准测试10款现成LLMs,发现其需求质量评估存在高漏报、误报及非单调性能等缺陷,暂不可作为自主评估工具,仅可作为人在环的决策支持。

AI 中文摘要

需求工程(RE)管控着系统工程(SE)下游所有事物的质量;未在评审环节被剔除的缺陷需求会蔓延至设计返工、进度延误和成本超支。由于需求常以自然语言编写,生成式AI的近期进展让人们对大语言模型(LLMs)能承担需求质量评估任务抱有期待,该任务原本耗时且依赖专业人力。然而,关于LLMs是否可信赖地完成该任务的实证证据仍十分匮乏。本研究针对现成LLMs在需求质量评估中的表现开展了首个基准分析。基于INCOSE质量准则构建的专家推导真值,我们评估了10个模型,涵盖OpenAI和Anthropic两大系列,每个系列各五代模型,评估过程涉及100次独立运行、两组需求集以及五种采样温度。研究得出四项贡献:其一,我们量化了高度不对称的误差特征:在所有模型和运行中,表现最佳的Anthropic模型仅检测到专家识别问题的中位数47%,同时误报率为11%;其二,在需要SE判断的场景中性能显著下降,必要性和正确性问题几乎总是被漏报;其三,代际进步并非单调,因此不能假设更新的模型表现更好;其四,这种误差行为在采样温度下仅发生小幅且非单调的变化,表明误差源于模型固有的缺陷而非内在随机性。因此,现成LLMs目前还不是可信赖的自主评估工具。研究结果也提醒Agentic AI开发者:在专用架构中编排这些LLM模块可能会加剧缺陷而非修正它们,其在近期的合理角色是作为人在环中的决策支持工具。

英文摘要

Requirements engineering (RE) governs the quality of everything downstream in systems engineering (SE); defective requirements that survive review cycles propagate into design rework, schedule delays, and cost overruns. Because requirements are often written in natural language, recent advances in generative AI have raised expectations that large language models (LLMs) can absorb requirement quality assessment, a task otherwise slow and human expertise-intensive. Yet empirical evidence on whether LLMs can be trusted to do so remains scarce. This study presents the first benchmarking analysis of off-the-shelf LLM performance for requirement quality evaluation. Against an expert-derived ground truth built on INCOSE quality criteria, we evaluate ten models spanning two families (OpenAI and Anthropic) and five generations each, across one hundred independent runs, two requirement sets, and five sampling temperatures. Four contributions follow. First, we quantify a strongly asymmetric error profile: across all models and runs, the best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11%. Second, performance degrades significantly where SE judgment is required, as necessity and correctness issues are almost always missed. Third, generational progress is non-monotonic, so newer models cannot be assumed better. Fourth, this error behavior shifts only modestly and non-monotonically across sampling temperatures, indicating characteristic model deficiencies rather than inherent stochasticity. Off-the-shelf LLMs are therefore not yet trustworthy autonomous evaluators. Findings also warrant caution for Agentic AI developers: orchestrating these LLM modules in specialized architectures risks compounding these deficiencies rather than correcting them. Their defensible near-term role is human-in-the-loop decision support.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑