发表机构
Development Data Group; Office of the World Bank Group Chief Statistician; The World Bank; World Bank UNHCR Joint Data Center(发展数据局; 世界银行集团首席统计学家办公室; 世界银行; 世界银行-联合国难民署联合数据中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种弱监督框架,利用LLM精化标签,从强迫流离失所与FCV文献中提取数据集提及,在1,706段落基准上实现74.1%精确率和70.5%召回率,为数据使用分析提供基础。
AI 中文摘要
发展和人道主义组织制作并支持调查、行政登记册及其他数据资源,以服务于研究、政策和运营,但系统地识别这些数据集被引用的位置仍然困难。这些引用分散在研究论文、项目文件、人道主义报告和其他非结构化文本中,限制了追踪数据使用以及识别数据可用性或传播方面潜在空白的能力。我们提出了一种弱监督框架,用于在不首先构建大型人工标注训练语料库的情况下,将数据集提取适配到强迫流离失所和脆弱、冲突与暴力(FCV)文献中。一个在通用研究文献上训练的轻量级模型从未标注的领域文档中生成候选数据集提及,前沿大语言模型(LLM)在上下文中审查这些候选,验证或拒绝它们并纠正其提取边界。生成的标注辅以有针对性的合成和对比示例,并用于微调轻量级模型以进行大规模提取。我们在一个独立的金标准基准上评估所得模型,该基准包含跨越研究、人道主义和运营文档的1,706个文本段落。在整个基准上,模型在提及级别达到74.1%的精确率和70.5%的召回率;在包含数据集引用的段落中,精确率达到89.5%。在段落级别,模型在区分包含数据集引用的段落与不包含的段落时达到88.2%的准确率和88.6%的特异性。这些结果展示了一种在标注数据有限时构建领域特定监督的实用方法,并为更大规模分析流离失所数据格局中的数据使用和潜在空白提供了技术基础。
英文摘要
Development and humanitarian organizations produce and support surveys, administrative registries, and other data resources to inform research, policy, and operations, yet systematically identifying where these datasets are referenced remains difficult. Such references are dispersed across research papers, project documents, humanitarian reports, and other unstructured text, limiting both the ability to trace data use and to identify potential gaps in data availability or dissemination. We present a weakly supervised framework for adapting dataset extraction to forced displacement and Fragile, Conflict, and Violence (FCV) documents without first constructing a large manually labeled training corpus. A lightweight model trained on general research literature generates candidate dataset mentions from unlabeled domain documents, which a frontier large language model (LLM) reviews in context, validating or rejecting candidates and correcting their extraction boundaries. The resulting annotations are supplemented with targeted synthetic and contrastive examples and used to fine-tune the lightweight model for large-scale extraction. We evaluate the resulting model on an independent gold-standard benchmark of 1,706 text passages spanning research, humanitarian, and operational documents. Across the full benchmark, the model achieves 74.1\% precision and 70.5\% recall at the mention level; among passages containing dataset references, precision reaches 89.5\%. At the passage level, the model achieves 88.2\% accuracy and 88.6\% specificity in distinguishing passages with dataset references from those without them. These results demonstrate a practical approach for constructing domain-specific supervision when labeled data are limited, and provide a technical foundation for larger-scale analysis of data use and potential gaps in the displacement data landscape.
Comments22 pages, 1 figure