发表机构
Harbin Engineering University(哈尔滨工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现SciCode基准存在大量缺陷,导致语言模型科学编码能力被低估,修正后重新评估显示模型准确率大幅提升,瓶颈在于评估工具而非模型能力。
AI 中文摘要
SciCode是衡量语言模型科学编码能力的标准指标,它针对的是需要前沿科学理论并将其实现为可运行数值代码的研究级问题,是《人工智能分析指数》的组成部分,也是政府及国家实验室套件中的常设评估项目。然而其分数近期陷入停滞:2026年最强模型的子问题准确率集中在60%左右,一款继任模型与前代表现持平。我们将这种停滞追溯到基准本身的缺陷:对全部65个测试问题按问题进行领域专家审计,发现263个缺陷,其中192个分布在91%的主要问题中,这些缺陷会导致符合指令的正确解决方案被错误拒绝,原因包括不可复现的标准答案、过严的容差或自相矛盾的规范;关键的是,78%的这类分数抑制型缺陷需要专业物理或数学知识才能检测,而非仅靠文书校对。我们修正了所有可确认的缺陷,生成SciCode-Verified,修正内容仅补充了良定问题所需的规范、修复了评分机制并收紧了过松的测试,每一项变更都附带理由并由第二位领域专家独立复核。我们在修正后的基准上重新评估了12款前沿模型快照,发现准确率大幅回升:子问题准确率从45%-60%升至84%-98%,主要问题准确率从9%-27%升至69%-92%。研究表明,最先进模型的科学编码能力远强于SciCode所显示的水平,瓶颈并非模型能力,而是评估工具的质量;我们发布SciCode-Verified及其完整审计轨迹作为修正后的公开标准。
英文摘要
SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code. It is a component of the Artificial Analysis Intelligence Index and a standing evaluation in government and national-laboratory suites. Yet its scores have recently plateaued: the strongest 2026 models cluster tightly around 60\% subproblem accuracy, and a successor model ties its predecessor. We trace this stagnation to defects in the benchmark itself. A per-problem, domain-expert audit of all 65 test problems uncovers 263 defects; 192 of them, spread across 91\% of the main problems, cause correct, instruction-following solutions to be wrongly rejected---through non-reproducible gold answers, over-tight tolerances, or self-contradictory specifications. Critically, 78\% of these score-suppressing defects require specialized physics or mathematics knowledge to detect, not mere clerical proofreading. We corrected every confirmable defect to produce SciCode-Verified. The corrections add only the specifications a well-posed problem requires, repair grading, and tighten the tests that were too lenient; every change is recorded with its justification and independently re-checked by a second domain expert. We re-evaluate twelve frontier model snapshots on the corrected benchmark and find a substantial recovery: subproblem accuracy rises from 45--60\% to 84--98\%, and main-problem accuracy from 9--27\% to 69--92\%. State-of-the-art models are far more proficient in scientific coding than SciCode has suggested---the bottleneck was not model capability, but the quality of the evaluation instrument. We release SciCode-Verified with its complete audit trail as the corrected public standard.
Comments47 pages, 2 figures, 6 tables. Project repository: https://github.com/flyingwagner/scicode-verified