使用基于大语言模型的实体解析在混合格式居住数据中进行家庭移动检测
Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution
浏览论文内容
中文总结 AI 辅助
研究在混合格式居住数据中检测家庭移动相关间接实体链接的问题,核心方法是集成基于提示的LLM命名实体识别、语义文本嵌入和基于图的推理,主要贡献是提高召回率和F1分数。
中文摘要 AI 辅助
实体解析(ER)通常依赖于记录之间的成对相似性比较,这限制了其捕捉人口居住数据中间接关系的能力。家庭移动会产生一种重要的间接模式,即多个人一起跨地址搬迁,但由于混合格式记录、噪声、重复以及缺乏稳定标识符,检测这种模式很困难。本文提出了一个人工智能增强框架,用于在未标准化的姓名-地址数据中检测与家庭移动相关的间接实体链接。该方法集成了基于提示的大语言模型(LLM)命名实体识别、语义文本嵌入和基于图的推理。对使用合成居住生成器生成的SPX基准数据集(S8-S12)的实验评估表明,纳入间接家庭移动证据可将召回率提高8-15%,同时保持高精度,F1分数比强大的成对基线提高6-8%。
英文摘要
Entity resolution (ER) typically relies on pairwise similarity comparisons between records, which limits its ability to capture indirect relationships present in demographic occupancy data. An important indirect pattern arises from household movement, where multiple individuals relocate together across addresses, but detecting such patterns is difficult due to mixed-format records, noise, duplication, and the absence of stable identifiers. This paper proposes an AI-enhanced framework for detecting indirect entity links associated with household movement in unstandardized name-address data. The approach integrates prompt-based large language model (LLM) named entity recognition for extracting personal names and addresses without extensive preprocessing, semantic text embeddings for robust similarity computation, and graph-based reasoning to infer group-level movement patterns. Experimental evaluation on SPX benchmark datasets (S8-S12) generated using the Synthetic Occupancy Generator demonstrates that incorporating indirect household movement evidence improves recall by 8-15% while maintaining high precision, yielding F1-score gains of 6-8% over a strong pairwise baseline.