AI 中文总结
本研究提出生成后融合策略,利用信息检索数据融合算子集成大语言模型的多次随机运行,以提升网络安全需求生成的召回率和精确率,实验表明均匀融合和朴素贝叶斯融合显著优于单次配置,将噪声转化为实用资产。
AI 中文摘要
将安全标准中的高层控制转化为具体、系统特定的需求,是网络安全需求工程的核心。大语言模型(LLMs)可以加速这一劳动密集且对召回敏感的任务,但任何单次运行都是不可靠的:它可能遗漏有效的保障措施,同时引入看似合理的幻觉,并且输出在不同运行和模型之间会发生变化。我们将这种变异性重新视为一种资源:我们不选择单一输出,而是研究生成后集成,使用信息检索数据融合算子聚合随机运行的结果。我们提出了两种策略:均匀融合仅奖励跨运行的一致性,而朴素贝叶斯融合则根据每个运行的估计可靠性进行加权。我们在来自四个模型家族的12种配置的24次运行上评估了这两种策略,这些运行针对十个ISO/IEC 27002:2022控制项生成,并由专家根据包含72条有效需求的金标准进行评判。汇集每次运行的输出可以恢复全部72条需求(而单一配置平均恢复不到一半),但也会产生111条幻觉。融合将麦粒与谷壳分开,将有效需求排在幻觉之前。在精确率-召回率和ROC曲线下面积方面,仅均匀融合就大幅超越了所有原始运行和配置,比最佳配置高出0.142和0.118。朴素贝叶斯加权进一步增加了0.039和0.052,达到0.864和0.869,同时更早地获得了有用的操作点。内部验证证实了这些收益的稳定性:在至少92%的袋外自举重采样和每个结构化扰动样本中,收益保持为正。因此,生成后融合将明显的噪声转化为实用资产:一个轻量级层,为分析师提供更广泛的覆盖范围和更优的优先级审查队列,仅使用经济实惠的、低于前沿的模型。
英文摘要
Translating high-level controls from security standards into concrete, system-specific requirements is central to cybersecurity requirements engineering. Large language models (LLMs) can accelerate this labor-intensive, recall-sensitive task, but any single run is unreliable: it misses valid safeguards while introducing plausible hallucinations, and outputs shift across runs and models. We reframe this variability as a resource: rather than selecting one output, we study post-generation ensembling, aggregating stochastic runs with information-retrieval data-fusion operators. We propose two strategies: Uniform fusion rewards mere cross-run agreement, whereas Naive-Bayes fusion weights each run by its estimated reliability. We evaluate both over 24 runs from 12 configurations across four model families, generated for ten ISO/IEC 27002:2022 controls and expert-judged against a gold standard of 72 valid requirements. Pooling every run's output recovers all 72 (whereas single configurations recover on average under half) but also 111 hallucinations. Fusion separates the wheat from the chaff, ranking valid requirements well ahead of hallucinations. In the areas under the precision-recall and ROC curves, Uniform fusion alone largely surpasses every original run and configuration by 0.142 and 0.118 over the best configuration. Naive-Bayes weighting adds a further 0.039 and 0.052, reaching 0.864 and 0.869 while attaining useful operating points earlier. Internal validation confirms the stability of these gains: they stay positive in at least 92% of out-of-bag bootstrap resamples and every structured-perturbation sample. Post-generation fusion thus turns apparent noise into a practical asset: a lightweight layer giving analysts broader coverage and a better prioritized review queue, using affordable, below-frontier models alone.
Comments20 pages; 4 figures; frozen candidate corpus, ensembling code, and computed artifacts (metrics, curves, result tables) archived on Zenodo, see https://doi.org/10.5281/zenodo.21496481