AI 中文总结
研究劳动力市场中职业编码问题,提出两步法,第一步用特定领域NER模型识别职业头衔,第二步映射到分类法,提高了准确性、鲁棒性和可解释性,还引入新置信标准,发布代码和脚本以支持可重复性。
AI 中文摘要
职业编码将自由文本中的职位头衔与职业分类法相联系,是劳动力市场研究的核心任务。现有方法通常在单个端到端步骤中解决此问题,同时识别职位头衔并分配职业代码。本文提出一种新颖的两步法来分离这些任务。第一步,特定领域的命名实体识别(NER)模型识别连续文本中的职业头衔,即使存在如OCR错误等噪声。第二步,将提取的职位头衔映射到分类法,使分类器专注于这种映射。实验证明,与单步方法相比,这种分离提高了准确性、鲁棒性和可解释性。该方法已针对德语文档开发,但可移植到其他语言。还引入了基于边际的职业编码置信标准,取代常见的绝对阈值。为支持可重复性,发布了源代码和评估脚本。
英文摘要
Occupation coding links job titles in free text to occupational taxonomies and is a core task in labor market research. Existing approaches typically address this problem in a single end-to-end step, jointly identifying job titles and assigning occupational codes. This paper presents a novel two-step approach that separates these tasks. In the first step, a domain-specific Named Entity Recognition (NER) model identifies occupational titles in continuous text, even under noise such as OCR errors. In the second step, the extracted job titles are mapped to a taxonomy, enabling the classifier to focus exclusively on this mapping. We demonstrate that this separation improves accuracy, robustness, and interpretability compared to single-step approaches. The method has been developed for German documents but is transferable to other languages. We further introduce a margin-based confidence criterion for occupation coding, replacing common absolute thresholds. To support reproducibility, we publish the source code and evaluation scripts.
CommentsPreprint of the paper accepted for the Federated Conference on Computer Science and Information Systems (FedCSIS 2026)
Journal refProceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS), IEEE, 2026