AI 中文总结
CreateScore提出基于领域理论的贝叶斯路由,用8B模型本地解决低不确定性简历筛选决策,升级高不确定性至120B模型,降低成本65.2%,但不确定性信号未能识别错误决策。
AI 中文摘要
大型语言模型(LLMs)可以支持基于评分标准的简历筛选,但对每个候选人和每条标准都应用高能力模型成本高昂。我们提出CreateScore,一个基于领域理论的贝叶斯网络,用于标准级别的LLM路由。一个手工指定的有向无环图,配备Dirichlet-multinomial条件概率表,将简历证据转换为后验不确定性;低不确定性的决策由本地8B模型解决,而不确定的决策则升级到120B参考模型。该图具有因果动机,但系统执行的是标准贝叶斯条件化,而非因果推断。升级阈值在训练折上校准(目标:70%本地解决),然后固定。在200份合成数据科学简历(139名训练和61名测试候选人,五个标准)上,77.7%的标准决策在本地解决(305项中的237项)。相对于120B模型裁决每条标准的参考条件,路由升级将令牌使用量减少了65.2%,并将精确分数一致性从32.8%(仅8B)提高到42.6%(95%置信区间31.0-55.1%);在n=61时,增益在统计上不显著。然而,不确定性信号并未识别出8B模型出错的决策:升级决策与参考的不一致率为16.2%,而本地解决决策为19.4%(AUROC 0.47,95%置信区间0.39-0.56),不比随机选择更好。我们还记录了早期评估如何因截断的推理模型输出被静默替换为本地标签而失效,并建议为级联评估提供保障措施。CreateScore被支持为一种可审计的成本降低机制,尚不是针对性的错误检测器,也不是自主招聘系统。
英文摘要
Large language models (LLMs) can support rubric-based screening of CVs, but applying a high-capability model to every candidate and criterion is costly. We present CreateScore, a domain-theory-informed Bayesian network for criterion-level LLM routing. A hand-specified directed acyclic graph with Dirichlet-multinomial conditional probability tables converts CV evidence into posterior uncertainty; low-uncertainty decisions are resolved by a local 8B model and uncertain ones are escalated to a 120B reference model. The graph is causally motivated, but the system performs standard Bayesian conditioning, not causal inference. The escalation threshold is calibrated on a training fold (target: 70% resolved locally) and then frozen. On 200 synthetic Data Science CVs (139 training and 61 test candidates, five criteria), 77.7% of criterion decisions were resolved locally (237 of 305). Relative to a reference condition in which the 120B model adjudicated every criterion, routed escalation reduced token use by 65.2% and raised exact score agreement from 32.8% (8B alone) to 42.6% (95% CI 31.0-55.1%); at n = 61 the gain was not statistically distinguishable. The uncertainty signal did not, however, identify the decisions on which the 8B model erred: disagreement with the reference was 16.2% among escalated and 19.4% among locally resolved decisions (AUROC 0.47, 95% CI 0.39-0.56), no better than random selection. We also document how an earlier evaluation was invalidated when truncated reasoning-model outputs were silently replaced by local labels, and we recommend safeguards for cascade evaluation. CreateScore is supported as an auditable cost-reduction mechanism, not yet as a targeted error detector, and is not an autonomous hiring system.
Comments11 pages, 5 figures, 7 tables (Excluding Appendix). CreateScore Planner App GitHub repo link: https://github.com/rupsa-arc1-2441139/createscore