发表机构
Carnegie Mellon University; University of Madison(卡内基梅隆大学; 麦迪逊大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出结合多维项目反应理论模型与问题上下文的LLM评估框架,可预测模型在未见问题上的表现,场景内评估效果优于无模型基线,但跨场景泛化仍为核心挑战。
AI 中文摘要
对大语言模型(LLM)的评估越来越需要在收集大量新标注前预测模型在新问题或任务上的表现,该问题颇具挑战性,因为问题难度、场景及底层能力需求会存在显著差异。简单的回顾性平均值可能混淆模型能力与项目特征。本文研究一种基于模型的评估框架,将多维项目反应理论模型与问题上下文相结合,以预测LLM在未见问题上的表现。该框架通过潜在能力画像表示LLM,同时利用问题内容明确项目特征,使信息能在已观测项目之外转移。实证研究发现,在场景内评估中,相较于无模型基线,融入问题嵌入可提升预测效果,且多维潜在结构比单维替代方案能更丰富地描述能力差异。同时,研究结果也揭示了一个重要局限:泛化性未必能转化为跨场景迁移下的可靠预测。这些发现表明,感知上下文的心理测量建模是高效且可解释的LLM评估的有前景方向,同时也凸显跨场景泛化是核心开放挑战。
英文摘要
Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.