发表机构
Malt(Malt)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对HR平台的多语言非标准化技能声明问题,提出结合LLM与Wikidata KG的智能体式混合KG生成流水线,可生成可扩展可解释的技能KG。
AI 中文摘要
组织数千个非标准化、多语言的专业技能声明是人力资源(HR)平台面临的长期挑战,直接影响人才精准匹配等下游任务。为解决该问题,我们提出一种混合知识图谱生成流水线,将大语言模型(LLM)与Wikidata多语言知识图谱(KG)结合,同时采用智能体反思模式合成新兴概念及其关联元数据。与僵化的自顶向下方法或碎片化的自底向上方法不同,我们的系统将已识别概念锚定到稳定的知识图谱实体,同时为未识别技能动态创建新节点和关系元数据。该系统分五个阶段执行:实体对齐、多语言规范化、主动筛选、去重及未映射概念的迭代恢复,可自主适应五种欧洲语言中快速演变、含噪声的技能提及。最终,此流水线提供了一个高度可扩展、可解释且具备自修复能力的框架,用于从非结构化、含噪声文本生成全面的技能知识图谱,并从中提取结构化分类体系。
英文摘要
Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybrid knowledge graph generation pipeline that grounds a Large Language Model (LLM) in the Wikidata multilingual Knowledge Graph (KG) while employing an agentic reflexion pattern to synthesize emerging concepts and their associated metadata. Unlike rigid top-down methods or fragmented bottom-up approaches, our system anchors recognized concepts to stable Knowledge Graph entities while dynamically creating new nodes and relational metadata for unrecognized skills. Executed across five stages, entity reconciliation, multilingual canonicalization, active curation, deduplication, and the iterative recovery of unmapped concepts, the system autonomously adapts to rapidly evolving, noisy skill mentions across five European languages. Ultimately, this pipeline provides a highly scalable, explicable, and self-healing framework for generating a comprehensive skills knowledge graph, from which a structured taxonomy is derived, using unstructured, noisy text.