arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAWILE:用于检查LLM评估器的多轴工作台

MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

Jackson Hassell, Farima Fatahi Bayat, Pouya Pezeshkpour, Estevam Hruschka

arXiv 2609.22599首次发表:更新:

发表机构

Megagon Labs(Megagon Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MAWILE是一个开发者工作台,通过四个表面(提示、标准、输入、输出)的受控扰动,审计LLM评估器的敏感性,无需黄金标签,可衡量鲁棒性和敏感性。

AI 中文摘要

大型语言模型(LLM)评估器为评估模型和智能体输出提供了一种灵活且可扩展的方法,但其裁决可能对评估响应中的偶然变化、评估器指令和评分标准敏感。现有系统检查了这些失败模式的重要子集,但审计一个配置好的评估器需要同时测试评估器工具及其评估的项目。我们引入了MAWILE,一个面向开发者的工作台,用于审计评估器在四个方面的敏感性:评估器提示、评估器标准、目标系统输入和目标系统输出。给定用户提供的评估器和代表性的评估项目,MAWILE构建并验证受控扰动,重新执行评估器,并定位由此产生的敏感性。每个扰动声明裁决应保持不变还是按指定方向变化,从而允许同一系统同时衡量对无关变化的鲁棒性和对有意义变化的敏感性。MAWILE审计二元、序数和成对评估器,无需黄金标签。该工具的代码可在以下网址获取:this http URL。

英文摘要

Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation items, MAWILE constructs and validates controlled perturbations, re-executes the judge, and localizes the resulting sensitivity. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes. MAWILE audits binary, ordinal, and pairwise judges without requiring gold labels. The code for this tool is available at: github.com/megagonlabs/mawile-judge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑