arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

G-SHARE:一种基于指南的人为因素事件诊断结构化推理框架

G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis

Xingyu Xiao, Mao Du, Jiejuan Tong, Jingang Liang, Haitao Wang

arXiv 2607.11892首次发表:更新:

发表机构

Institute of Nuclear and New Energy Technology, Tsinghua University; National Key Laboratory of Human Factors Engineering; Fujian Fuqing Nuclear Power Co., Ltd.(清华大学核能与新能源技术研究院; 人因工程重点实验室; 福建福清核电有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对核电站人为因素事件诊断问题,提出基于指南的结构化推理框架G-SHARE,将诊断指南转化为多阶段流程,含证据提取等步骤。通过构建数据集评估,其性能显著优于基线,证明了转化专家指南为推理流程对安全关键行业智能分析的价值。

AI 中文摘要

人为因素事件诊断对于从核电站运行事件中学习至关重要,但其质量很大程度上依赖专家对叙述性报告的解读。基于指南的数据驱动或一次性大语言模型方法往往缺乏结构化推理,与正式诊断指南的一致性有限,且可能产生逻辑不一致的结论。本研究提出G-SHARE,一个基于指南的结构化推理框架,将CNNP九步人为因素事件诊断指南转化为多阶段诊断流程。该框架由证据提取﹑逐步诊断推理和事后一致性修复组成,能明确使用报告证据、生成中间理由并对诊断输出进行逻辑验证。通过来自中国核工业源的真实人为因素事件报告构建数据集,并使用领域专家注释的金标准子集进行评估。结果表明,G-SHARE显著优于一次性提示和传统机器学习基线,最强版本实现了最佳总体准确率和宏F1。消融结果进一步表明,结构化推理和一致性执行对于稳健诊断至关重要,特别是在弱提示条件下。研究结果证明了将专家诊断指南转化为可审计推理工作流程的价值,为安全关键行业的智能人为因素分析提供了实用途径。

英文摘要

Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert interpretation of narrative reports and guideline-based reasoning.Existing data-driven or one-shot large language model approaches often lack structured reasoning, have limited alignment with formal diagnostic guidelines, and may generate logically inconsistent conclusions. To address this issue, this study proposes G-SHARE, a guideline-based structured reasoning framework that operationalizes the CNNP nine-step human-factor event diagnosis guideline into a multi-stage diagnostic pipeline.The framework consists of evidence extraction, stepwise diagnostic reasoning, and post-hoc consistency repair, enabling explicit use of report evidence, intermediate rationale generation, and logical validation of diagnostic outputs. A dataset of real human-factor event reports was constructed from Chinese nuclear industry sources, and a gold-standard subset annotated by domain experts was used for evaluation. Results show that G-SHARE substantially outperforms one-shot prompting and traditional machine learning baselines, with the strongest version achieving the best overall accuracy and macro-F1. Ablation results further indicate that structured reasoning and consistency enforcement are critical to robust diagnosis, especially under weak prompting conditions. The findings demonstrate the value of transforming expert diagnostic guidelines into auditable reasoning workflows, providing a practical pathway for intelligent human-factor analysis in safety-critical industries.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑