LLJ Cards:大型语言模型作为评判者的最佳实践
LLJ Cards: Best practices for the Use of LLMs as Judges
浏览论文内容
中文总结 AI 辅助
本文提出LLJ Cards框架,综合测量理论、自然语言生成和机器学习文献的最佳实践,为基于大型语言模型作为评判者的评估提供标准化、透明且可复现的实用指南,以确保评估的有效性、可靠性和可复现性。
中文摘要 AI 辅助
近年来,大型语言模型(LLMs)已成为评估领域的一种流行替代方案。这些系统通常被称为“大型语言模型作为评判者”(LLJs),已被研究人员和从业者广泛采用于各种测量任务,这得益于其相对于人类判断的强性能、可扩展性和成本效益。然而,越来越多的研究表明,使用LLJs作为评估者引发了对其有效性和可靠性的担忧。现有应对这些挑战的努力主要集中在开发偏差缓解技术和优化提示策略上。虽然这些方法代表了重要的一步,但它们主要提供技术性修复,却留下了一个更根本的挑战未得到解决:缺乏标准化、透明且可复现的评估实践。在本文中,我们引入了LLJ Cards,这是一个框架,将测量理论、自然语言生成和机器学习文献中的最佳实践综合为基于LLJ的评估的实用指南。虽然LLJs为可扩展评估提供了一条有前景的路径,但其有效使用需要以严格的评估原则为基础,以确保有效性、可靠性和可复现性。LLJ Cards通过提供一个结构化框架来满足这一需求,该框架用于在自动化评估的设计和报告中应用这些原则。
英文摘要
In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-effectiveness relative to human judgment. However, a growing body of work has shown that the use of LLJs raise concerns about their validity and reliability as evaluators. Existing efforts to address these challenges have largely focused on developing bias-mitigation techniques and refining prompting strategies. While these approaches represent an important step forward, they primarily offer technical fixes and leave a more fundamental challenge unaddressed: the lack of standardized, transparent, and reproducible evaluation practices. In this paper, we introduce LLJ Cards, a framework that synthesizes best practices from measurement theory, natural language generation, and machine learning literature into practical guidelines for LLJ-based evaluations. While LLJs offer a promising path toward scalable evaluation, their effective use requires grounding in rigorous evaluation principles to ensure validity, reliability, and reproducibility. LLJ Cards addresses this need by providing a structured framework for applying these principles in the design and reporting of automated evaluations.
发表机构
- McGill University(麦吉尔大学)
- Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
- Concordia University(康考迪亚大学)
机构由 AI 辅助整理,请以论文原文为准。