AI 中文总结
CalibratedRubric是结合特定类型评分、贝叶斯过滤与IRT的任务自适应框架,可提升开放式LLM评估的人工-黄金标准一致性与排序保真度,减少所需评分规则数量。
AI 中文摘要
对开放式大语言模型(LLM)输出的可靠评估需要细粒度评分规则,但专家编纂成本高昂且难以规模化。现有自动化流程依赖严格的评判者一致性和二元方差过滤,无法区分可测量评分规则与信息性评分规则。我们提出CalibratedRubric,这是一种任务自适应框架,结合了特定类型评分、贝叶斯评分规则可测量性过滤以及基于项目反应理论(IRT)的规则库组装。CalibratedRubric通过Beta-伯努利一致性后验估计每个评分规则的可测量性,并采用子模块信息覆盖目标,在观测能力范围内构建紧凑的评分规则库。在金融、医疗、通用和法律基准测试中,可测量性过滤使JudgmentBench上的人工-黄金标准一致性从κ=0.604提升至0.743;基于IRT的贪婪选择在所有6个评估的响应块上,相较于随机选择提高了交叉拟合的排序保真度,且在FinResearchBench决策支持任务中仅需49个而非131个评分规则即可达到目标相关性;任务标签扰动进一步降低了系统分离度,证实了任务自适应评分的实际相关性。这些结果表明CalibratedRubric是一种高效、感知不确定性的开放式LLM评估方法,其校准增益取决于足够的评判者冗余度。
英文摘要
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $κ=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.