勿宣称面向基准的优化可提升通用编码能力——需要多样化评估
Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required
浏览论文内容
中文总结 AI 辅助
该研究指出少量编码基准的优化无法提升通用编码能力,通过案例研究验证了基准排名难泛化、跨任务迁移差等问题,呼吁采用差异化评估与持续基准维护。
中文摘要 AI 辅助
后训练论文、模型卡片和博客文章常将少量编码基准(如SWE-bench和LiveCodeBench)上的分数作为研究成果和面向用户系统具备广泛编码能力的证据。我们认为,针对这些基准的优化会导致测量任务特定性能,在测得分数与通用编码能力的宣称间产生意义鸿沟。我们通过创建的基于Django的案例研究基准套件考察该鸿沟,评估在SWE-bench轨迹上后训练的基础模型和检查点,发现基准排名常无法泛化;后训练检查点几乎无跨任务迁移,SWE-bench优化在我们的任务或LiveCodeBench上仅产生有限或无增益,同样,对单个Django模态的微调也无法迁移。我们得出结论,少量基准不足以在基准优化压力下评估多样化模型,鼓励社区采用差异化评估:前沿模型的整体评估、研究用多任务套件、窄任务应用的人在环研究,还主张创建能力分类体系和持续的基准维护,而非一次性基准发布。若无可靠评估标准,使用LLMs和智能体的工程师与研究者将不得不依赖不充分证据做出研究、开发和部署决策。
英文摘要
Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.g., SWE-bench and LiveCodeBench) as evidence of broad coding capability, both for research artifacts and user-facing systems. We argue that optimization for these benchmarks leads to measuring task-specific performance, creating a meaning gap between measured scores and claims of general coding ability. We examine this gap with a Django-based case study benchmark suite we create. Evaluating foundation models and checkpoints post-trained on SWE-bench trajectories, we find that benchmark rankings frequently fail to generalize. Post-trained checkpoints show little cross-task transfer, and SWE-bench optimization yields limited or no gains on our tasks or on LiveCodeBench. Similarly, fine-tuning on individual Django modalities fails to transfer. We conclude that a small number of benchmarks is insufficient for evaluating diverse models under benchmark optimization pressure. We encourage the community to use differentiated evaluation - holistic assessment for frontier models, multi-task suites for research, and human-in-the-loop studies for narrow task applications. Finally, we argue for creating a capability taxonomy and sustained benchmark maintenance, rather than one-off benchmark releases. Without reliable evaluation standards, engineers and researchers using LLMs and agents have to rely on insufficient evidence to make research, development, and deployment decisions.