发表机构
BroutonLab; Curately(BroutonLab; Curately)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出通过蒸馏生产级LLM信号构建可解释的简历-职位匹配特征表示,包含基于LLM的离线标注器和CPU可运行的蒸馏双编码器,在168,772对数据上训练,达到95.79%的招聘人员一致性。
AI 中文摘要
将候选人与职位进行匹配是招聘的核心环节,而招聘人员不仅需要看到一个单一的、不透明的相关性分数,还需要了解候选人为何匹配。我们以命名的、可解释的匹配维度形式提供这一证据——在当前部署中包括八个维度。我们提出了一种由两部分组成的方法。第一部分是基于LLM的标注器,其提示和特征定义在作为早期生产匹配阶段时,根据招聘人员的反馈进行了优化。在当前架构中,它仅用于离线标注,不参与在线请求。第二部分是从中蒸馏出的特征双编码器:一个经过LoRA适配的嵌入骨干网络,配备紧凑的逐维度头,可在CPU上运行并处理所有在线请求。两部分都在持续改进:随着反馈的到来修订提示,双编码器则根据更新的标签进行重新训练。该模型在168,772对带标签的职位-简历对(17,921个职位和180,030份简历)上进行了训练。使用该服务的招聘人员可以确认或修改呈现出的特征预测。在来自该选定生产反馈子集的927个招聘人员记录值中,部署的学生模型在888个案例(95.79%)中与记录决策一致。这是操作性的、非盲态的一致性,而非独立的人工评估。
英文摘要
Matching candidates to vacancies is central to recruitment, and a recruiter needs to see why a candidate fits, not only a single opaque relevance score. We provide this evidence as named, interpretable matching dimensions recruiters can act on - eight in our current deployment. We propose a two-part approach. The first is an LLM-based labeler whose prompts and feature definitions were refined from recruiter feedback while it served as an earlier production matching stage. In the current architecture, it is used only for offline labeling and is not called on online requests. The second is a feature bi-encoder distilled from it: a LoRA-adapted embedding backbone with compact per-dimension heads that runs on CPU and serves all online requests. Both parts keep improving: prompts are revised as feedback arrives, and the bi-encoder is retrained on the updated labels. The model is trained on 168,772 labeled vacancy-resume pairs (17,921 vacancies and 180,030 resumes). Recruiters using the service can confirm or revise surfaced feature predictions. On 927 recruiter-recorded values from this selected production-feedback subset, the deployed student agrees with the recorded decisions in 888 cases (95.79%). This is operational, non-blinded agreement rather than an independent human evaluation.
CommentsAccepted to EMNLP 2026; 13 pages, 4 figures, 7 tables