JobHop v2:一个来自非结构化简历的大规模职业轨迹数据集
JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes
- AIDA-IDLab, Department of Electronics and Information Systems, Ghent University, Ghent, Belgium(AIDA-ID实验室,电子与信息系统系,根特大学,根特,比利时)
- AIDA-IDLab, Department of Electronics(AIDA-ID实验室,电子系)
- Information Systems, Ghent University, Ghent, Belgium(信息系统,根特大学,根特,比利时)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出JobHop v2数据集,通过端到端大语言模型从大量简历中提取职业轨迹。它扩大了覆盖范围与注释丰富度,引入新提取管道、模式及评估协议,最好提取器接近注释者间一致性上限,数据集和代码公开助力职业轨迹研究。
AI中文摘要:
大规模、注释丰富的职业轨迹数据对劳动力规划、工作推荐和劳动力市场分析至关重要,但公开可用的数据集要么规模小、不便于独立使用,要么基于预标准化职业代码构建,使用大语言模型合成而非真实自由文本。我们展示了JobHop v2,它是公开可用的JobHop数据集的改进版本,通过对弗拉芒公共就业服务机构VDAB提供的约440,000份化名多语言简历语料库进行端到端大语言模型提取构建而成。发布的数据集包含355,315条职业轨迹,用ESCO职业代码、季度级时间信息和标准化五级教育程度进行注释,扩大了原始版本的覆盖范围和注释丰富度。相对于v1,JobHop v2引入了基于推理控制的大语言模型推理和重试机制的重新设计提取管道(实现100% JSON解析率)、更丰富的提取模式以及针对三个互补注释基线评分的修订评估协议。在与这些基线的评估中,我们最好的提取器在所有比较模型中最接近注释者间一致性上限,仅落后1.1 - 2.7个百分点。数据集和代码已公开发布以支持可重复的职业轨迹研究。
英文摘要:
Large-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from pre-standardized occupational codes with LLM-synthesized rather than authentic free text. We present JobHop~v2, an improved version of the publicly available JobHop dataset, constructed through end-to-end large language model (LLM) extraction from a corpus of ${\sim}440{,}000$ pseudonymized, multilingual resumes provided by VDAB, the Flemish Public Employment Service. The released dataset comprises $355{,}315$ career trajectories annotated with ESCO occupational codes, quarter-level temporal information, and normalized five-level education attainment, broadening both the coverage and the annotation richness of the original release. Relative to v1, JobHop~v2 introduces a redesigned extraction pipeline based on reasoning-controlled LLM inference with a retry mechanism (achieving a 100% JSON parse rate), a richer extraction schema, and a revised evaluation protocol scored against three complementary annotation baselines. Evaluated against these baselines, our best extractor comes closest to the inter-annotator agreement ceiling among all compared models, trailing it by only 1.1-2.7 percentage points. The dataset and code are publicly released to support reproducible career-trajectory research.