发表机构
AIDA-IDLab, Ghent University(AIDA-ID实验室,根特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用大语言模型从简历获取职业轨迹数据,提出STEP职业路径推荐系统,通过集成时间衰减GRU、FiLM及注意力序列池化预测下一份工作,引入ROUTE改进职业表示,在多数据集评估中优于基准,还公开了数据集和代码。
AI 中文摘要
职业路径蕴含着数十年的技能获取、角色转换和教育投资,大规模理解职业路径对劳动力规划、劳动力市场政策和工作推荐至关重要。简历是职业路径信息的丰富来源,但因其非结构化、异构和多语言的性质,长期阻碍大规模系统分析。随着大语言模型的出现,现在可以从非结构化简历中获取包含时间和教育信号的丰富职业轨迹数据,为职业路径推荐带来新机遇。利用这一机会,我们提出了STEP(就业预测顺序轨迹),这是一种新颖的职业路径推荐系统,它利用时间和教育信号来预测职业轨迹中的下一份工作。STEP集成了一个时间衰减门控循环单元(GRU)来建模时间动态,基于教育程度的特征线性调制(FiLM),以及基于注意力的序列池化来选择用于下一份工作预测的相关特征。为了改进STEP的内部职业表示,我们引入了ROUTE,这是一个两阶段对比过程,首先通过无监督去噪自编码器使多语言编码器适应职业领域,然后通过引导负选择进行监督对比微调。我们在四个职业轨迹数据集上评估了STEP,包括我们公开可用的JobHop数据集的改进版本,并表明它在下一份工作预测方面优于现有基准。数据集和代码已公开发布,以支持可重复的职业轨迹研究。
英文摘要
Career paths encode decades of skill acquisition, role transitions, and educational investment, and understanding them at scale underpins workforce planning, labor market policy, and job recommendation. Resumes are a rich source of information about career paths: they contain detailed descriptions of work experience, education, and skills. Yet their unstructured, heterogeneous, and multilingual nature has long prevented large-scale systematic analysis. With the advent of large language models (LLMs), it is now possible to source rich career trajectory data containing temporal and educational signals from unstructured resumes, enabling new opportunities for career-path recommendation. Exploiting this opportunity, we present STEP (Sequential Trajectory of Employment Prediction), a novel career-path recommendation system that leverages temporal and educational signals to predict the next job in a career trajectory. STEP integrates a time-decay Gated Recurrent Unit (GRU) cell to model temporal dynamics, Feature-wise Linear Modulation (FiLM) conditioned on educational attainment, and attention-based sequence pooling to select relevant features for next job prediction. To improve internal occupation representation for STEP, we introduce ROUTE, a two-stage contrastive procedure that first adapts a multilingual encoder to the career domain via unsupervised denoising autoencoding, then performs supervised contrastive fine-tuning with guided negative selection. We evaluate STEP on four datasets of career trajectories, including an improved version of our publicly available JobHop dataset, and show that it outperforms state-of-the-art baselines in next job prediction. The dataset and code are publicly released to support reproducible career-trajectory research.