arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37322cs.AIcs.SE

Mubric:面向LLM评估的基于变异测试的评分标准生成

Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation

Jiayuxuan Yang, Jie M. Zhang, Yiling Lou, Zhenpeng Chen

首次发表
浏览论文内容

中文总结 AI 辅助

Mubric提出基于变异测试的评分标准生成方法,通过挖掘缺陷并注入参考响应来迭代改进评分标准,在703个任务上以7.48个百分点的优势超越最强基线。

中文摘要 AI 辅助

基于评分标准的评估广泛用于评估基于LLM的系统,其做法是将响应质量分解为特定任务的评分标准。然而,自动生成能够可靠捕获特定任务质量要求的评分标准仍然具有挑战性。我们提出Mubric,一种基于变异测试的评分标准生成方法。变异测试是一种经典的软件测试方法,通过向程序中注入故障并检查测试是否检测到这些故障来评估测试套件。我们将测试套件与评分标准进行类比:如果评分标准捕获了重要的质量要求,那么向原本高质量的响应中引入相应缺陷应降低其得分。Mubric首先从偏好响应与非偏好响应的真实配对中挖掘常见缺陷,并将这些缺陷抽象为可复用的变异算子,每个算子指定如何引入特定类型的响应缺陷。对于新任务,它将相关算子应用于参考响应,检查注入的缺陷是否降低响应质量,并利用惩罚不足的缺陷来改进评分标准。我们在四个代表性领域的703个任务上,将Mubric与六种先进的评分标准生成方法进行比较。Mubric取得了最高的整体评估准确率,比最强基线高出7.48个百分点。

英文摘要

Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.

发表机构

  • Tsinghua University(清华大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑