arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31275cs.SE

软件决策支持的贝叶斯网络半自动量化:在软件研发组织中比较WSA与RNM

Semi-Automatic Quantification of Bayesian Networks for Software Decision Support: Comparing WSA and RNM in a Software R&D Organization

  • Federal University of Campina Grande(坎皮纳格兰德联邦大学)
  • Aarhus University(奥胡斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Mirko Perkusich, João Nunes, Emilia Mendes, Emanuel Dantas, Ademar Sousa, Danyllo Albuquerque, Kyller C. Gorgônio, Angelo Perkusich

AI总结:

本研究通过嵌入式案例比较WSA与RNM两种半自动量化方法,发现它们虽在决策结果上相似,但概率分布差异显著,强调方法选择需考虑节点语义、专家判断和输出分布等因素。

AI中文摘要:

当历史数据有限时,专家驱动的贝叶斯网络可以支持重复出现的软件决策,但量化其条件概率表(CPTs)需要大量概率判断。半自动方法减少了直接启发,但从业者缺乏来自实际决策过程的比较性证据。我们报告了一个嵌入式案例研究,在单个软件研发(R&D)组织内的两个情境中,比较了加权和算法(WSA)和排名节点方法(RNM)作为完整的启发与量化流程:物联网项目的特征选择以及用户界面设计选择。在每个情境中,流程共享图结构、根先验和决策证据。我们通过15个专家定义的模型走查场景和决策会议中记录的备选方案的回顾性重构来评估它们。WSA匹配了7/7和6/8的走查预期,而RNM匹配了4/7和4/8。两种流程都将选定的备选方案置于回顾性排名的顶部附近。然而,它们的完整分布不同:四个不同特征模式的中位数总变差距离为0.304,七个不同设计模式为0.340,其中一个设计模式升至0.890。这些结果涉及一个组织中的两个序数价值估计模型。它们表明,相似的短名单行为并不等同于对不确定性的等效表示。对于可比较的软件决策支持模型,方法选择和验证应考虑节点语义、专家能提供的判断、校准要求、完整输出分布以及概率的预期下游用途。

英文摘要:

Expert-driven Bayesian networks can support recurring software decisions when historical data are limited, but quantifying their conditional probability tables (CPTs) requires many probability judgments. Semi-automatic methods reduce direct elicitation, yet practitioners have little comparative evidence from operational decision processes. We report an embedded case study comparing the Weighted Sum Algorithm (WSA) and Ranked Nodes Method (RNM) as complete elicitation-and-quantification pipelines in two contexts within a single software research and development (R&D) organization: feature selection for Internet of Things projects and user interface design selection. Within each context, the pipelines shared the graph, root priors, and decision evidence. We evaluated them through 15 expert-defined model-walkthrough scenarios and retrospective reconstructions of alternatives recorded in decision meetings. WSA matched 7/7 and 6/8 walkthrough expectations, whereas RNM matched 4/7 and 4/8. Both pipelines placed the selected alternatives near the top of the retrospective rankings. Their complete distributions nevertheless differed: the median total variation distance was 0.304 across four distinct feature patterns and 0.340 across seven distinct design patterns, rising to 0.890 for one design pattern. These results concern two ordinal value-estimation models in one organization. They show that similar shortlist behavior does not imply equivalent representations of uncertainty. For comparable software decision-support models, method selection and validation should consider node semantics, the judgments experts can provide, calibration requirements, complete output distributions, and the intended downstream use of the probabilities.

补充信息

↑