不可作弊评估:基于动态压缩的语言模型评估
Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models
浏览论文内容
中文总结 AI 辅助
提出不可作弊评估(Uncheatable Eval),通过动态收集新文本并利用压缩率评估基础语言模型,降低数据污染风险;实验表明压缩性能随模型规模缩放,且与零样本MMLU准确率负相关。
中文摘要 AI 辅助
现代大型语言模型在海量数据集上进行预训练,这使得防止基准数据进入其训练集并破坏评估结果的可靠性变得困难。对于基础模型而言,可靠的评估尤其具有挑战性,因为其有限的指令遵循能力使基于任务的评估变得复杂。我们引入了不可作弊评估(Uncheatable Eval),这是一个动态基准,定期收集新发表的文本以评估基础语言模型并降低数据污染的风险。利用模型预测能力与其无损压缩数据能力之间的关系,我们使用压缩率来评估模型预测新文本的效果。我们评估了80个模型在14个文本类别上的表现,研究了压缩率如何随上下文长度变化,并考察了压缩率与零样本MMLU准确率之间的相关性。我们的结果得出三个主要发现:(1)压缩性能随模型规模呈现一致的缩放趋势;(2)基于注意力、混合和循环模型在更多上下文可用时压缩性能的变化方式不同;(3)较低的压缩率与较高的零样本MMLU准确率密切相关。代码可在 https://github.com/Jellyfish042/uncheatable_eval 获取。
英文摘要
Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to evaluate base language models and reduce the risk of data contamination. Drawing on the relationship between a model's predictive ability and its ability to compress data losslessly, we use compression rate to evaluate how well models predict new text. We evaluate 80 models across 14 text categories, study how compression changes with context length, and examine the correlation between compression rate and zero-shot MMLU accuracy. Our results yield three main findings: (1) compression performance follows a consistent scaling trend with model size; (2) attention-based, hybrid, and recurrent models differ in how their compression performance changes as more context becomes available; and (3) lower compression rates are strongly associated with higher zero-shot MMLU accuracy. Code is available at https://github.com/Jellyfish042/uncheatable_eval.
发表机构
- Shenzhen University(深圳大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。