arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HazardWeaver:面向灾害分析智能体的科学路线选择

HazardWeaver: Scientific Route Selection for Hazard Analysis Agents

Wangshu Zhu, Xueqi Cheng, Liang Wu, Yushun Dong

arXiv 2610.03591首次发表:更新:

发表机构

Florida State University; Nokia(佛罗里达州立大学; 诺基亚)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HazardWeaver提出状态依赖的科学路线选择方法,通过知识编译和能力图,为灾害分析智能体选择并调整可执行的科学路线,在141个实例基准上优于现有系统。

AI 中文摘要

理解和评估自然灾害对于灾害防备和风险降低至关重要。大型语言模型的最新进展激发了人们对用于灾害分析的AI智能体的日益增长的兴趣,特别是它们将科学数据、模型和工具整合到自动化工作流程中的能力。然而,有效的自动化要求智能体确定哪些科学方法适用于给定事件,并且能够利用现有数据和工具执行。随着新的证据和执行结果的出现,这些条件可能发生变化,要求智能体重新考虑其选择。我们将此问题形式化为状态依赖的科学路线选择,并引入HazardWeaver。具体而言,HazardWeaver首先利用Hazard知识编译器提取与证据关联的、控制科学适用性的条件,然后其Hazard能力图表示可执行的科学能力并检查其输入和输出之间的兼容性。利用这些互补的表示,HazardWeaver智能体组件选择适用且可执行的路线,执行其工作流程,并在分析状态变化时修订其决策。为了评估科学输出以及产生这些输出的决策,我们引入了HazardWeaver基准,包含跨越七个单一灾害领域和四个多灾害交互类别的141个实例。该基准容纳多条有效的科学路线,并评估输出正确性、路线有效性以及合理的弃权(不执行)。在该基准上的大量实验表明,HazardWeaver优于现有的智能体系统,在具有多条合格科学路线的任务上取得了最大的提升。我们的代码在此https URL公开可用。

英文摘要

Understanding and assessing natural hazards is essential for disaster preparedness and risk reduction. Recent advances in large language models have spurred growing interest in AI agents for hazard analysis, particularly their ability to integrate scientific data, models, and tools into automated workflows. However, effective automation requires agents to determine which scientific methods are appropriate for a given event and executable with the available data and tools. As new evidence and execution results become available, these conditions can change, requiring agents to reconsider their choices. We formulate this problem as state-dependent scientific route selection and introduce HazardWeaver. Specifically, HazardWeaver first leverages the Hazard Knowledge Compiler to extract evidence-linked conditions governing scientific applicability, then its Hazard Capability Graph represents executable scientific capabilities and checks compatibility between their inputs and outputs. Using these complementary representations, the Hazard Weaver Agent component selects applicable and executable routes, carries out their workflows, and revises its decisions as the analysis state changes. To evaluate both the scientific outputs and the decisions that produce them, we introduce the Hazard Weaver Benchmark, comprising 141 instances across seven single-hazard domains and four multi-hazard interaction classes. The benchmark accommodates multiple valid scientific routes and evaluates output correctness, route validity, and justified abstention. Extensive experiments on this benchmark show that HazardWeaver outperforms existing agent systems, with the largest gains on tasks with multiple eligible scientific routes. Our code is publicly available at https://github.com/LabRAI/HazardWeaver.

Comments24 pages, including references and appendices. Code is available at https://github.com/LabRAI/HazardWeaver

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑