AI 中文总结
本文提出单令牌期望值评分法,通过微调小型语言模型实现冷启动候选排名,在求职者和雇主相关性上均显著优于基线,并提升在线性能。
AI 中文摘要
AI辅助采购简化了候选人审查流程,减轻了招聘人员手动筛选的行政负担。然而,将语言模型部署为生产级排名器仍然具有挑战性。零样本大型语言模型(LLMs)可能产生不稳定、非确定性的分数,排名准确性较低,而传统深度神经排名器需要数百万条记录交互,这是低流量、小众采购平台无法产生的。相反,可用的只有几十万个序数相关性标签——按排名器训练标准来看规模较小,但当预训练语言模型已经编码了任务依赖的一般世界知识时,这些标签就足够了。我们提出了单令牌期望值评分,这是一种排名原语,将候选人-职位相关性视为对等级令牌{1,...,5}的序数分类,并将相关性分数读取为第一令牌概率分布的期望值。由于分数来自单次解码步骤而非开放式生成,它是模型logits的确定性函数,无需输出解析,且延迟低。为了仅从这种监督中学习异构招聘标准的非线性相互依赖关系,我们使用混合序数回归损失微调了一个小型语言模型(SLM),该损失结合了均方误差项(保留序数距离)和分类交叉熵项(锐化类别边界)。我们沿着两个维度——求职者相关性和雇主相关性——使用NDCG@10和低相关性率进行评估。离线状态下,我们微调的模型优于启发式基线和零样本LLMs。端到端模拟显示了相同方向且幅度更大(求职者NDCG@10提高54.2%,低相关性率降低46.7%),实时在线实验将雇主低相关性降低了27.3%,并将雇主保留率提高了7.07%。
英文摘要
AI-assisted sourcing streamlines candidate review, reducing the administrative burden of manual screening for recruiters. However, deploying language models as production rankers remains challenging. Zero-shot Large Language Models (LLMs) may produce unstable, non-deterministic scores and rank less accurately, while conventional deep neural rankers require millions of logged interactions that a low-traffic, niche sourcing platform does not produce. What is available instead is a few hundred thousand ordinal relevance labels -- small by ranker-training standards, but sufficient when a pretrained language model already encodes the general world knowledge the task depends on. We present single-token expected-value scoring, a ranking primitive that casts candidate-job relevance as an ordinal classification over the grade tokens {1, ..., 5} and reads the relevance score as the expectation of the first-token probability distribution. Because the score comes from a single decoding step rather than open-ended generation, it is a deterministic function of the model's logits, requires no output parsing, and serves at low latency. To learn the non-linear interdependencies of heterogeneous hiring criteria from this supervision alone, we fine-tune a Small Language Model (SLM) with a hybrid ordinal regression loss combining a Mean Squared Error term, which preserves ordinal distance, with a categorical Cross-Entropy term, which sharpens class boundaries. We evaluate along two dimensions -- Jobseeker Relevance and Employer Relevance -- using NDCG@10 and low relevance rate. Offline, our fine-tuned model outperforms a heuristic baseline and zero-shot LLMs. An end-to-end simulation shows the same direction at larger magnitude (+54.2% Jobseeker NDCG@10, -46.7% low relevance rate), and a live online experiment reduces employer low-relevance by 27.3% and raises employer keep rate by 7.07%.
Comments10 pages, 7 figures. Accepted at RecSys in HR '26: The 6th Workshop on Recommender Systems for Human Resources, in conjunction with the 20th ACM Conference on Recommender Systems (RecSys 2026), September 28 - October 2, 2026, Minneapolis, MN, USA. To appear in CEUR Workshop Proceedings