arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

泛化即稳定性,而非准确性:大语言模型的多轴评估

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi

arXiv 2610.01428首次发表:更新:

发表机构

Technion – Israel Institute of Technology(以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出稳定性感知泛化目标(SAGO),从多维度评估LLM泛化稳定性,发现常用模型存在显著且一致的泛化不稳定性,且跨数据集变化可逆转模型排名。

AI 中文摘要

大语言模型(LLM)中的泛化能力是指当同一输入以不同方式表达时,模型能够产生一致且语义稳定的输出。现有工作通常通过在单一提示格式、任务或一组变体上的聚合准确性来评估泛化,这混淆了鲁棒性与整体基准性能。在本工作中,我们展示了在单个示例层面、跨多种输入变体以及跨模型行为的不同方面进行泛化评估,重点关注变异性,而不是将性能简化为一个可通过狭窄训练或其他掩盖泛化评估的方式提高的分数。遵循这一观点,我们引入了稳定性感知泛化目标(SAGO),这是一个衡量模型行为在相同输入的不同变体和基准下变化程度的框架,捕获了包括生成一致性、内部激活、置信度和响应镜像在内的多个维度的变异性。我们表明,许多常用模型表现出统计上显著且一致的泛化不稳定性:没有模型能均匀地泛化,行为轴捕获了独立的失败模式,且跨数据集的变化可以逆转模型排名。

英文摘要

Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robustness with overall benchmark performance. In this work, we show generalization evaluation at the level of individual examples, across multiple input variants, and across different aspects of model behavior, focusing on variability rather than reducing performance to a score that can be improved through narrow training or other ways that obfuscate generalization evaluation. Following this view, we introduce the Stability-Aware Generalization Objective (SAGO), a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring. We show that many commonly used models exhibit statistically significant and consistent generalization instability: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.

CommentsAccepted at the TAE (Trust-AI-Eval) Workshop: Can We Trust AI Evaluation?, NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑