arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从匹配模型到招聘智能体:AI招聘系统、评估与治理的系统化叙事综述

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

Ziyi Zhao, Guanzheng Wei

arXiv 2609.04286首次发表:更新:

发表机构

University of Chinese Academy of Social Sciences; Southwest University(中国社会科学院大学; 西南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述梳理AI招聘系统从匹配模型到招聘智能体的发展,分析三大转变,指出领域缺口并提出评估映射与系统议程,明确进展评判标准。

AI 中文摘要

人工智能在招聘领域的应用已将自动化对象从简历配对与排名列表转变为多阶段工作流,用于检索证据、比较候选人并支持或执行相关行动。本系统化叙事综述追溯了该领域的发展历程,从双向检索与行为排名,到神经人岗匹配、大语言模型(LLM)组件,再到使用工具的招聘智能体。本研究采用目的性检索与编码方案,截至2026年7月23日已完成更新,并于2026年9月2日进行了针对性更新,共梳理了40篇代表性研究及相关产业与法律资料。本综述并非患病率估计,而是分析了三个相互关联的转变:从相似性到双向适配,从单一模型到复合工作流,从离线预测到与证据及生产力对齐的评估。在文档理解、检索、排名、评估、面试、人才寻访与人工交接等环节,本研究区分了领域级、配对级、列表级、案例级、轨迹级与结果级证据。研究发现存在持续的缺口:行为标签混淆了曝光度、偏好与任职资格;私有数据与合成数据限制了外部有效性;最终输出分数掩盖了流程缺陷;在编码的研究中,隐私未被直接评估,且没有任何一项研究同时评估效用、公平性、隐私与安全性。这些观察结果仅描述了编码的研究集合,而非整个领域。因此,本研究引入了从评估证据到最强可辩护主张的阶段性映射,并提出了一个关于双向、基于证据、时间可控、选择性及可审计系统的研究议程。进展的评判标准应是工作流是否能检索到正确证据、保留不确定性、支持可争议决策,并在明确的成本与风险约束下改善结果。

英文摘要

Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person--job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation. Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security. These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.

Comments53 pages, 4 figures, 10 tables; companion literature-coding and search-log CSVs included in the source package

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑