How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
大型语言模型在评估中作弊了多少?基于一次性密码本的框架下的高估基准测试
机构 * Tech Startups(科技初创公司)
专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL
AI总结 针对大型语言模型在公开基准测试中因数据污染或训练偏差导致评估结果虚高的问题,提出基于一次性密码本加密思想的动态评估框架ArxivRoll,包含自动生成私有测试用例的SCP模块和衡量污染与偏差比例的Rugged Scores指标,实现可重复、透明且高效的评估。
Comments This paper has been accepted by AAAI 2026. We update it for adding new evaluation results for ArxivRollBench-2025a and ArxivRollBench-2026a, with the evaluation of timly models like DeepSeekV4Pro, GPT-5.5, Claude-Opus-4.7, and so on. Source code: https://github.com/liangzid/ArxivRoll/ Online Leaderboard Website: https://arxivroll.moreoverai.com/