arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SESSE:草图、扩展、排序、总结、评估——通过结构化分解的大模型作为评判者的评估

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

Dae Lee, Mihai Delgeanu, Adel Youssef

arXiv 2608.18303首次发表:更新:

发表机构

Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大模型作为评判者评估无法分离质量维度等问题,提出无需训练的SESSE框架,在RewardBench上实现与基线近乎相当的性能,且可提供可解释的审计跟踪。

AI 中文摘要

大模型作为评判者的评估将响应质量评估简化为单一的整体A/B偏好选择,无法分离出驱动偏好的质量维度,也无法区分模型错误与真实的标签歧义。我们提出SESSE(Sketch, Expand, Sort, Summarize, Evaluate),这是一种无需训练的框架,它将整体判断分解为从评判者自身错误案例中直接挖掘的结构化子问题;无需参考响应、特定任务的评分标准或微调。在RewardBench(n=1000)上,SESSE实现了与思维链基线近乎相当的性能,且与微调的专用模型RISE-Judge-32B(92.7%)具有竞争力,同时保持完全无需训练。每个标准的投票证据提供了可解释的审计跟踪,用于诊断单一整体输出令牌无法提供的标签歧义与评判者失败模式。

英文摘要

LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into structured sub-questions mined directly from the judge's own error cases; requiring no oracle responses, task-specific rubrics, or fine-tuning. On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of-thought baseline and is competitive with RISE-Judge-32B (92.7%), a fine-tuned specialist, while remaining fully training-free. Per-criterion vote evidence provides an interpretable audit trail for diagnosing label ambiguity and judge failure modes unavailable from a single holistic output token.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑