发表机构
University of Sindh; Mehran University of Engineering & Technology(信德大学; 迈赫兰工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出HUMAID-NER数据集及不确定性加权多任务学习框架,联合进行灾难推文命名实体识别与事件分类,在验证集上分别达到0.841和0.761的F1分数。
AI 中文摘要
从社交媒体中快速提取结构化信息对人道主义响应至关重要,然而现有的灾难推文资源主要提供文档级别的类别标签,缺乏跨度级别的实体标注。我们引入了HUMAID-NER,这是首个基于HumAID基准构建的命名实体识别数据集,包含60,000条英文灾难推文,以BIO格式标注了十个基于操作动机的实体类型,产生了约175,000个带标签的实体跨度。标注通过一个可复现的三阶段混合流程生成,该流程结合了spaCy transformer模型、灾难领域EntityRuler模式和结构化正则表达式,并采用基于优先级的重叠解决机制。我们还提出了一个联合多任务学习框架,使用共享的RoBERTa-large编码器同时进行灾难特定命名实体识别和人道主义事件分类。为减少联合训练中的任务冲突,模型采用同方差不确定性加权,使用可学习的任务参数,并采用两阶段训练计划,在第二阶段冻结24个编码器层中的底部18层。在HUMAID-NER验证集上,所提出的系统同时实现了NER跨度微F1分数0.841和分类宏F1分数0.761。一个实时网络仪表板展示了端到端的部署。数据集、模型和流程代码已发布以支持可复现性和未来的危机信息学研究。
英文摘要
Rapid extraction of structured information from social media is important for humanitarian response, yet existing disaster tweet resources mainly provide document-level category labels without span-level entity annotations. We introduce HUMAID-NER, the first named entity recognition dataset built on the HumAID benchmark, containing 60,000 English disaster tweets annotated in BIO format across ten operationally motivated entity types and yielding approximately 175,000 labelled entity spans. Annotations are generated through a reproducible three-stage hybrid pipeline combining a spaCy transformer model, disaster-domain EntityRuler patterns, and structured regular expressions with priority-based overlap resolution. We also propose a joint multitask learning framework that performs disaster-specific named entity recognition and humanitarian event classification using a shared RoBERTa-large encoder. To reduce task conflict during joint training, the model uses homoscedastic uncertainty weighting with learnable task parameters and a two-stage training schedule that freezes the lower 18 of 24 encoder layers in the second stage. On the HUMAID-NER validation set, the proposed system achieves NER span micro-F1 of 0.841 and classification macro-F1 of 0.761 simultaneously. A real-time web dashboard demonstrates end-to-end deployment. The dataset, models, and pipeline code are released to support reproducibility and future crisis informatics research.
Comments8 pages, 8 figures, 4 tables. Published in The Asian Bulletin of Big Data Management, Vol. 6, No. 1, pp. 138-152, 2026
Journal refThe Asian Bulletin of Big Data Management, 6(1), 138-152 (2026)