arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22559cs.AIcs.CLcs.IR

ExecRubrics:用于可验证且高效的长文本评估的可执行工具增强型评分标准

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein

首次发表
浏览论文内容

中文总结 AI 辅助

ExecRubrics是将评分标准转化为可执行Python程序的框架,可替代黑盒LLM评判者,在三个长文本基准上实现与自然语言评分标准相当或更优的偏好准确率,且评估延迟降低最多320倍,提升了评估的可解释性与效率。

中文摘要 AI 辅助

评分标准旨在通过将响应质量分解为可解释的标准,使语言模型评估更透明。然而,自然语言评分标准往往存在歧义,需要黑盒大语言模型(LLM)作为评判者,且通常假设标准通过线性加权和独立聚合,限制了其捕捉依赖关系、替代项、惩罚和覆盖条件的能力。我们提出ExecRubrics,一个将评分标准表示为紧凑可执行程序的框架。ExecRubrics将评估逻辑编码为可验证的Python评分函数,赋予自然语言评分标准意图操作语义:一种可检查、执行和编辑的固定决策程序。在三个长文本响应基准HealthBench、HelpSteer和ArgQuality上,我们表明ExecRubrics可替代昂贵的黑盒评判者来对偏好响应和非偏好响应进行排序,在最佳偏好准确率分别为53%、78%和92%的情况下,与自然语言评分标准基线相当或更优,同时将评估延迟降低多达320倍。我们还表明,整合来自NLTK和spaCy等文本处理库的外部逻辑和资源可进一步提高偏好准确率。我们的结果提出了一种看待评估的新方式,为黑盒评分标准评估提供了一种更快、更可解释且歧义更少的替代方案,尤其适用于医疗保健和银行等精度和可审计性至关重要的高风险领域。

英文摘要

Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require LLM judges, and typically assume criteria aggregated through linear weighted sums, limiting their ability to capture dependencies, alternatives, penalties, and override conditions. We propose ExecRubrics, a framework for representing rubrics as compact executable programs. ExecRubrics encodes evaluation logic as verifiable Python scoring functions, giving natural-language rubric intent an operational semantics: a fixed decision procedure that can be inspected, executed, and edited. On three long-form response benchmarks -- HealthBench, HelpSteer, and ArgQuality -- we show that ExecRubrics can recover substantial preference signal without an LLM judge at evaluation time. On ArgQuality and HelpSteer, the strongest executable variants are within 1.1 and 4 percentage points, respectively, of the direct GPT-5.5 agentic baseline. Executable rubrics are also considerably faster, achieving a 192x average speedup. We show that incorporating external logic and resources from text processing libraries such as NLTK and spaCy can further improve preference accuracy. Our results suggest a novel way of approaching automated evaluation, by offering a faster, more explainable, and less ambiguous alternative to black-box rubric evals, particularly in high-stakes domains such as healthcare and banking where precision and auditability are critical.

发表机构

  • Emory University(埃默里大学)
  • University of Waterloo(滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑