顺序正确,尺度错误:审计用于职业AI测量的LLM评判者
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
浏览论文内容
中文总结 AI 辅助
本研究通过O*NET-BENCH审计套件评估33种LLM评判者配置,发现排序一致性高但接受率估计偏差大,表明评判者需针对目标聚合指标验证。
中文摘要 AI 辅助
LLM评判者越来越多地被用于评估AI输出是否满足工作场所要求,但响应排序的一致性并不能确立接受率或职业聚合指标的一致性。我们引入了O*NET-BENCH,一个源自现有包含45,796条工人评分的调查的审计套件,并评估了六个模型家族中的33种预配置评判者配置,测试评分为4,501条。二十五种配置达到了至少0.60的平局感知对准确率,尽管一个训练拟合的仅响应TF-IDF基线几乎与最强评判者相当。尽管存在这种排序一致性,评判者估计3.0%-97.9%的响应是可接受的,而职业匹配的工人则为61.1%。在一个微调谱系中,从逐点评分改为捆绑式少样本/列表式协议改善了响应排序,同时降低了在任务和职业层面与工人均值的一致性;这种反转在预先指定标准下,在任务和工人不相交的验证分割上得到了复现。交叉验证校准在很大程度上消除了均值偏差,但校准后的分数最多只能解释个体工人评分方差的8.5%。在所研究的标签预算下,预测辅助估计最多带来微小的精度提升。这些结果表明,仅靠排序一致性不足以进行职业测量。评判者应针对其分数将被用于估计的接受率和聚合指标进行验证。
英文摘要
LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。