arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从UNDRR报告到事件记录:基于模式约束的LLM地理参考灾害抽取

From UNDRR Reports to Event Records: Schema-Constrained LLM Extraction of Georeferenced Disasters

Camilla Andreozzi, Phuong-Anh Nguyen-Le, Zhijing Jin, Revati Mani

arXiv 2609.23853首次发表:更新:

发表机构

ETH Zürich; University of Maryland; MPI; University of Toronto; United Nations Office for Disaster Risk Reduction(苏黎世联邦理工学院; 马里兰大学; 马克斯·普朗克研究所; 多伦多大学; 联合国减少灾害风险办公室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于模式约束的LLM流水线,从UNDRR的PreventionWeb文档中抽取地理参考灾害事件记录,在10,000份文档上生成3,572条记录,GPT-5达到86.0%的F1分数,显著优于基线,并计划开源以适配国家档案。

AI 中文摘要

减灾档案以散文形式描述灾害事件,而EM-DAT(Delforge等,2025)等数据库无法直接摄取这些内容。我们提出一个LLM流水线,使用受控灾害词汇表和固定模式生成候选地理参考事件记录,并保留证据以供审查。将该方法应用于UNDRR管理的知识中心PreventionWeb中的10,000份文档,生成了来自1,913份文档、涵盖24种灾害类型的3,572条记录,并将81%的地点提及解析为OpenStreetMap几何图形。在来自分层217份文档参考集的171个人工标注阳性文档窗口中,GPT-5实现了86.0%的合并属性F1分数,而spaCy-gazetteer基线为44.2%。评估按灾害族、地点字符串和事件年份对文档内进行汇总,但不评估它们对单个事件的分配。GPT-5.4在十个LLM中排名最高(86.6% F1)。逐字证据出现率方面,GPT-5为72.0%,GPT-5.4为47.2%,这衡量了文本可追溯性,但未确立属性支持。我们报告了生产中的失败模式以及自动化标签和位置规则合规性检查。提示词、模式和输出将发布,以便适应国家报告档案。

英文摘要

Disaster-risk-reduction archives describe hazard events in prose that databases such as EM-DAT (Delforge et al., 2025) cannot ingest directly. We present an LLM pipeline that generates candidate georeferenced event records using a controlled hazard vocabulary and fixed schema, retaining evidence for review. Applied to 10,000 documents from PreventionWeb, the knowledge hub managed by UNDRR, it produced 3,572 records from 1,913 documents across 24 hazard types and resolved 81% of location mentions to OpenStreetMap geometries. On 171 human-positive document windows from a stratified 217-document reference set, GPT-5 achieved 86.0% pooled attribute $F_1$, versus 44.2% for the spaCy-gazetteer baseline. Evaluation pools hazard families, location strings, and event years within documents, without assessing their assignment to individual events. GPT-5.4 ranked highest among ten LLMs (86.6% $F_1$). Verbatim evidence occurrence was 72.0% for GPT-5 and 47.2% for GPT-5.4, measuring textual traceability without establishing attribute support. We report production failure modes and automated label and location-rule compliance checks. Prompts, schema, and outputs will be released for adaptation to national reporting archives.

Comments17 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑