arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越分数:理解摘要评估中的大语言模型作为评判者的机制

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

Himil Vasava, Ming Jiang

arXiv 2609.01604首次发表:更新:

发表机构

University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM作为评判者的摘要评估机制,通过多组实验分析了Themis和Prometheus等模型的评估流水线结构及微调的作用,并发布了相关代码与数据。

AI 中文摘要

基于大语言模型(LLM)的自然语言生成(NLG)质量评估器被广泛用作评分工具和自动训练信号,但它们分配评分的内部过程仍未被充分理解。我们通过针对NLG质量的可读性和充分性维度的八项攻击扰动分类法、生成具有可控错误强度和显式token级修改映射的成对干净与损坏摘要的生成流水线,以及应用于Themis(Llama-3-8B)和Prometheus(Mistral-7B)的因果追踪、logit-lens词汇投影和注意力头 knockout 的四项实验,从机制上研究了这一过程。两种评估器都实现了一个结构化、连贯的两阶段评估流水线:在第15层以下,注意力执行局部错误比较并将结果路由到最终输入位置;在第15层以上,MLP级联整合信号并写入评分,决策在较晚的层(Themis为L=26,Prometheus为L=25)的残差流中明确形成。此外,相同规模的基础模型(Llama-3-8B)重现了路由架构和决策形成,但未重现阶段分离,从而分离出微调专门引入的两种机制:抑制第15层以下MLP在最后位置的贡献,以及决策形成深度提前两层,表明微调是在现有基础上塑造而非从头构建流水线。我们在该https URL发布了源代码和数据。

英文摘要

LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and explicit token-level modification maps, and a four-experiment battery of causal tracing, logit-lens vocabulary projection, and attention-head knockout applied to Themis (Llama-3-8B) and Prometheus (Mistral-7B). Both evaluators implement a structured, coherent evaluation pipeline operating in two stages: below layer 15, attention performs local error comparison and routes the result to the final input position; above it, the MLP cascade integrates the signal and writes the rating, with the decision crystallizing in the residual stream at a sharp late layer (L = 26 on Themis, L = 25 on Prometheus). Furthermore, a base-model control at the same scale (Llama-3-8B) reproduces the routing architecture and crystallization but not the stage separation, isolating the two mechanisms that fine-tuning specifically installs, suppression of below-L15 MLP contribution at the last position and a two-layer advance of the crystallization depth, indicating that fine-tuning sculpts an existing substrate rather than building the pipeline from scratch. We release the source code and data at https://github.com/himil-v/judge-mech

CommentsAccepted at EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑