arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于余弦相似度重构的变分自编码器从航空事故叙述中推断语义因果因素

Semantic Causal-Factor Inference from Aviation Incident Narratives Using A Variational Autoencoder with Cosine-Similarity-Based Reconstruction

AZIIDA NANYONGA1, HASSAN WASSWA, UGUR TURHAN, KEITH FRANCIS JOINER, GRAHAM WILD

arXiv 2610.04472首次发表:更新:

发表机构

SDAIA-KFUPM Joint Research Center for Artificial Intelligence, King Fahd University of Petroleum and Minerals (KFUPM); University of New South Wales(SDAIA-KFUPM人工智能联合研究中心,法赫德国王石油与矿产大学; 新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出结合自然语言处理与变分自编码器的框架,从航空事故叙述中推断语义因果因素,在2万余份NTSB报告上实现0.786的余弦相似度,展示了AI辅助提取因果信息的潜力。

AI 中文摘要

航空事故和事件调查会产生大量非结构化文本信息,其中包含与安全事件原因和促成因素相关的证据。自动提取此类信息具有挑战性,因为因果证据可能分布在冗长且复杂的调查叙述中。本研究提出了一种语义因果因素推断框架,该框架结合自然语言处理与变分自编码器(VAE),以学习航空调查叙述与专家报告的可能原因之间的关系。调查叙述及其对应的可能原因被转换为数值表示,随后编码器将叙述表示映射到概率潜在空间。解码器估计对应可能原因的表示,并使用结合Kullback-Leibler散度与基于余弦相似度的语义重构的目标进行训练。该框架使用2005年至2020年间20,919份已定稿的美国国家运输安全委员会调查报告中进行了评估。在保留的测试集上,预测表示与专家报告的可能原因表示达到了0.786的平均余弦相似度(标准差=0.120)。预测表示还产生了与报告中因果信息相关的可解释术语。结果表明,概率潜在表示学习在AI辅助从航空安全叙述中提取因果信息方面具有潜力,同时保留专家调查作为正式因果确定的基础。

英文摘要

Aviation accident and incident investigations generate extensive unstructured textual information containing evidence relevant to the causes and contributing factors of safety occurrences. Automatically extracting such information is challenging because causal evidence may be distributed across long and complex investigation narratives. This study proposes a semantic causal-factor inference framework combining natural language processing with a variational autoencoder (VAE) to learn the relationship between aviation investigation narratives and expert-reported probable causes. Investigation narratives and their corresponding probable causes are transformed into numerical representations, after which the encoder maps narrative representations to a probabilistic latent space. The decoder estimates representations of the corresponding probable causes and is trained using an objective that combines Kullback-Leibler divergence with cosine-similarity-based semantic reconstruction. The framework was evaluated using 20,919 finalized U.S. National Transportation Safety Board investigation reports from 2005 to 2020. On the held-out test set, the predicted and expert-reported probable-cause representations achieved a mean cosine similarity of 0.786 (SD = 0.120). The predicted representations also yielded interpretable terms associated with causal information in the reports. The results demonstrate the potential of probabilistic latent representation learning for AI-assisted extraction of causal information from aviation safety narratives while retaining expert investigation as the basis for formal causal determination.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑