AI 中文总结
该研究针对自然语言描述的人才挖掘任务,提出首个交互式需求驱动型候选人挖掘智能体,在21个系统和691项需求上,其广度和召回率均优于现有基线模型。
AI 中文摘要
从自然语言描述(如“转行到生物科技研究岗位的机器学习工程师”)中寻找人才的任务,越来越多地被委托给大型语言模型(LLM)智能体,并被框架化为信息检索任务。本文认为,该任务本质上是一个需求工程任务:此类请求是一个存在隐式约束的欠确定需求,存在大量有效答案且无明确验收标准,因此有用的答案需要在搜索前先引出、验证和确认需求。我们提出了\textit{\textbf{\textbackslash sys{}}}——据我们所知,首个交互式、需求驱动型候选人挖掘智能体(它通过有限引出、工作流模板、两阶段提交协议和双向终止防护,将模糊的人才请求转化为合理的候选名单,涵盖引出、验证、检索和确认环节),以及\textit{\textbf{\textbackslash bench{}}}——一个运行需求生命周期的基准(包括锚定标准的验证、多模型证据驱动的基准构建和成本感知的确认)。在21个系统和全部691项需求上,\textit{\textbf{\textbackslash sys{}}}在广度上占据主导地位(覆盖率达100%,产出量是其他系统的2.5倍),且与该领域“近正交”,其返回的人才中有90%是20个强大的LLM加网络基线模型都未发现的。除广度外,对每个系统的证据驱动评判显示,\textit{\textbf{\textbackslash sys{}}}召回了最多相关真实人才:在联合人才池中占0.241,是排名第二系统的1.9倍,且其自助法95%置信区间与所有基线模型均不重叠。因此,\textit{\textbf{\textbackslash sys{}}}是最强的挖掘引擎(拥有最深的可及真实候选人才池),而用于精度排序的LLM则作为互补的验证工具。
英文摘要
Finding people from a natural-language description (``ML engineers transitioning to research roles in biotech'') is increasingly delegated to LLM agents and framed as information retrieval. We argue that it is fundamentally a requirements engineering task: such a request is an under-determined requirement with implicit constraints, many valid answers, and no acceptance criterion, so useful answers require eliciting, validating, and verifying the requirement before search can matter. We present \sys{}, to our knowledge the first interactive, requirements-driven candidate-sourcing agent (it elicits, validates, retrieves, and verifies a vague people-request into a justified slate through bounded elicitation, workflow templates, a two-stage commit protocol, and bidirectional termination guards) and \bench{}, a benchmark that runs the requirements lifecycle (criteria-anchored validation, multi-model evidence-grounded oracle construction, and cost-aware verification). Across $21$ systems and all $691$ requirements, \sys{} dominates breadth ($100%$ coverage at $2.5\times$ the yield) and is \emph{near-orthogonal} to the field, with $90%$ of the people it returns are surfaced by \emph{none} of $20$ strong LLM-plus-web baselines combined. Beyond breadth, an evidence-grounded judging of every system shows \sys{} \emph{recalls} the most relevant real people: $0.241$ of the union pool, $1.9\times$ the next system, with a bootstrap $95%$ interval disjoint from every baseline. \sys{} is thus the strongest \emph{sourcing} engine (the deepest real, reachable candidate pool), while precision-ranking LLMs serve as~complementary verifiers.
Comments12 pages