arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18595cs.SE

OdinEval:用于基于大语言模型的Odin编程语言程序修复的可复现基准测试

OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language

Bang Xie, Hao Liu, Zhiyuan Peng, Xin Yin, Senjian Zhang, Yuan Luo, Chenhao Ying, Haiming Jin, Wei Chen, Shaocong Long, Zhenyu Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对Odin语言程序修复的空白,本文构建可复现基准OdinEval,评估6个语言模型,Kimi-K3解决率66.7%、Qwen3.8-Max复现率96.4%,发布配套多类资源。

中文摘要 AI 辅助

仓库级程序修复基准测试仍以少数主流语言为中心,Odin等系统语言基本未被测试。本文提出OdinEval,这是一个从公开Odin仓库中已记录缺陷构建的可复现基准测试。每个实例将问题与基础提交、修复提交、正确补丁、问题特定回归测试、历史工具链及执行记录绑定。准入要求为测试在基础修订版上失败,在正确修复后通过。当无可用开发者测试时,需通过同一模型的三个实例独立审查黑盒测试,在两种历史状态下执行,并根据带版本控制的测试编写技能的记录反馈进行修订。我们在168个筛选实例上,采用统一协议评估六个语言模型。Kimi-K3的已解决分数最高,为66.7%;Qwen3.8-Max的复现分数最高,为96.4%。发布内容包含冻结数据、源代码归档、容器、验证器、模型补丁及审计清单。

英文摘要

Repository-level repair benchmarks still center on a few mainstream languages, leaving systems languages such as Odin largely untested. We present OdinEval, a reproducible benchmark built from documented defects in public Odin repositories. Each instance binds an issue to base and fix commits, a gold patch, an issue-specific regression test, a historical toolchain, and execution records. Admission requires the test to fail on the base revision and pass after the gold fix. When no usable developer test exists, a black-box test is reviewed independently by three instances of the same model, executed in both historical states, and revised from recorded feedback under a versioned Test Writing Skill. We evaluate six language models on 168 filtered instances under one shared protocol. Kimi-K3 records the highest Resolved score at 66.7%, while Qwen3.8-Max has the highest Repro score at 96.4%. The release includes frozen data, source archives, containers, validators, model patches, and audit manifests.

补充信息

↑