arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过表示多任务学习进行LLM基准测试

LLM Benchmarking via Representation Multi-task Learning

Yuqing Xie, Yuxuan Xu, Yang Feng, Yunxiao Chen

arXiv 2610.05613首次发表:更新:

发表机构

London School of Economics and Political Science; New York University(伦敦政治经济学院; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于表示多任务学习和项目反应理论的统计框架,用于在排行榜中更准确地评估LLM的领域特定与综合能力,并通过模拟和MMLU数据验证其优越性。

AI 中文摘要

量化和评估大型语言模型(LLM)的能力仍然是现代数据科学和人工智能中的一个基本挑战。在本文中,我们考虑在排行榜框架内,基于LLM在多个基准领域(例如数学推理和编码)中的项目上的表现来进行评估。我们的目标是解决两个核心问题:(1)如何通过跨领域借用信息来获得更准确的领域特定分数?(2)如何定义和估计一个聚合多个领域表现的综合分数?为了解决这些问题,我们提出了一种基于表示多任务学习(MTL)和项目反应理论模型的新型统计框架。具体来说,我们通过项目反应理论(IRT)模型定义整体和领域特定的LLM特质,并提出一种MTL方法来从项目级响应数据中估计这些特质。我们开发了一种计算高效的估计器,并在某些渐近机制下建立了其极小极大最优性。该框架为系统性的LLM评估提供了严格的测量基础。我们进行了广泛的模拟,证明了所提出方法相对于竞争方法的优越性能。对于应用和案例研究部分至关重要的一点是,我们将所提出的框架应用于来自Hugging Face Open LLM排行榜的MMLU响应数据,涵盖了4,272个LLM和56个学科中的13,232个项目。实证分析揭示了领域规模和难度上的显著异质性,以及强烈的正跨领域依赖性,突出了我们方法所产生的实用价值和实质性见解。

英文摘要

Quantifying and evaluating the capabilities of Large Language Models (LLMs) remains a fundamental challenge in modern data science and artificial intelligence. In this paper, we consider LLM evaluation based on their performance across items in multiple benchmark domains (e.g., mathematical reasoning and coding) within a leaderboard framework. Our goal is to address two core questions: (1) How do we derive more accurate domain-specific scores by borrowing information across domains? and (2) How do we define and estimate an overall score that aggregates performance across multiple domains? To solve these problems, we propose a novel statistical framework based on representation Multi-task Learning (MTL) and an item response theory model. Specifically, we define overall and domain-specific LLM traits through an Item Response Theory (IRT) model, and propose an MTL approach to estimate these traits from item-level response data. We develop a computationally efficient estimator and establish its minimax optimality under certain asymptotic regimes. This framework provides a rigorous measurement foundation for systematic LLM evaluation. We conduct extensive simulations, demonstrating the superior performance of the proposed method over competing methods. Crucially for the Applications and Case Studies section, we apply the proposed framework to MMLU response data from the Hugging Face Open LLM Leaderboard, covering 4,272 LLMs and 13,232 items across 56 subjects. The empirical analysis reveals substantial heterogeneity in domain size and difficulty, together with strong positive cross-domain dependence, highlighting the practical value and substantive insights generated by our approach.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑