arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图结构评分标准:将评分标准编译为用于大语言模型评判器的类型化评估图

Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

Xi Chen, Jie Mu, Mo Xuan, Qun Shao

arXiv 2608.12097首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出图结构评分标准(GSR),将评分标准编译为类型化评估图,在GPT-OSS-120B上较Prometheus式评分提升了精确分数一致性,且在成对评估中取得更高准确率。

AI 中文摘要

基于评分标准的评估器通常将评分标准视为提示上下文或扁平准则:它们规定了要评判的内容,但即便自然语言规则已明确说明,准则的组合方式仍未明确。我们提出图结构评分标准(Graph-Structured Rubrics, GSR),该方法在观察响应前将评分标准编译为与响应无关的类型化评估图。准则节点引出评判结果;转换、归约和门控算子通过命名端口对其进行组合;而被称为“读出(Readout)”的特定任务输出映射则将唯一汇点转换为分数或偏好。编译过程会拒绝格式错误或类型不兼容的图。逐点评估会在图聚合前分别评判评分标准维度;成对评估则复用该图,对每个候选对象在每项准则下各进行一次评判。在GPT-OSS-120B模型下,GSR在四个逐点数据集上较Prometheus式评分将精确分数一致性提升了0.62至6.75个百分点,并在两个偏好基准上,采用原生平局与弃权(不执行)策略时取得了数值最高的端到端成对准确率。

英文摘要

Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compiles a rubric into a response-independent typed evaluation graph before observing responses. Criterion nodes elicit judgments; transformation, reduction, and gating operators compose them through named ports; and a task-specific output mapping, termed Readout, converts the unique sink into a score or preference. Compilation rejects malformed or type-incompatible graphs. Pointwise evaluation judges rubric dimensions separately before graph aggregation; pairwise evaluation reuses the graph with one judgment for each candidate under every criterion. Under GPT-OSS-120B, GSR improves exact score agreement by 0.62--6.75 percentage points over Prometheus-style scoring on four pointwise datasets and achieves the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.

Comments11 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑