OdinEval:用于基于大语言模型的Odin编程语言程序修复的可复现基准测试
OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language
浏览论文内容
中文总结 AI 辅助
针对Odin语言程序修复的空白,本文构建可复现基准OdinEval,评估6个语言模型,Kimi-K3解决率66.7%、Qwen3.8-Max复现率96.4%,发布配套多类资源。
中文摘要 AI 辅助
仓库级程序修复基准测试仍以少数主流语言为中心,Odin等系统语言基本未被测试。本文提出OdinEval,这是一个从公开Odin仓库中已记录缺陷构建的可复现基准测试。每个实例将问题与基础提交、修复提交、正确补丁、问题特定回归测试、历史工具链及执行记录绑定。准入要求为测试在基础修订版上失败,在正确修复后通过。当无可用开发者测试时,需通过同一模型的三个实例独立审查黑盒测试,在两种历史状态下执行,并根据带版本控制的测试编写技能的记录反馈进行修订。我们在168个筛选实例上,采用统一协议评估六个语言模型。Kimi-K3的已解决分数最高,为66.7%;Qwen3.8-Max的复现分数最高,为96.4%。发布内容包含冻结数据、源代码归档、容器、验证器、模型补丁及审计清单。
英文摘要
Repository-level repair benchmarks still center on a few mainstream languages, leaving systems languages such as Odin largely untested. We present OdinEval, a reproducible benchmark built from documented defects in public Odin repositories. Each instance binds an issue to base and fix commits, a gold patch, an issue-specific regression test, a historical toolchain, and execution records. Admission requires the test to fail on the base revision and pass after the gold fix. When no usable developer test exists, a black-box test is reviewed independently by three instances of the same model, executed in both historical states, and revised from recorded feedback under a versioned Test Writing Skill. We evaluate six language models on 168 filtered instances under one shared protocol. Kimi-K3 records the highest Resolved score at 66.7%, while Qwen3.8-Max has the highest Repro score at 96.4%. The release includes frozen data, source archives, containers, validators, model patches, and audit manifests.