arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06411cs.SEcs.AIcs.CL

RuBench:一个具有原生编写的俄语任务规范的仓库级智能编码基准测试

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

Evgeny Shilov

首次发表
浏览论文内容

中文总结 AI 辅助

RuBench 1.0是含25个俄语任务的仓库级智能编码基准测试,任务源于开源仓库修复提交,按客户请求风格编写。评估多种产品配置,报告运行结果,发现最佳配置解决78.7%任务,还发现产品替换模型问题,为智能编码评估提供新基准。

中文摘要 AI 辅助

开发者越来越多地将实际维护工作委托给产品级编码代理,许多任务用母语表述,类似客户请求而非精心策划的英文问题。现有仓库级智能基准测试未针对此设置:其任务声明设计为英文。我们引入RuBench 1.0,它包含从五个活跃开源仓库(aiohttp、aiogram、Laravel、NestJS、Fastify;Python、PHP、TypeScript、JavaScript)的近期修复提交中挖掘的25个任务,每个任务用俄语原生编写,按实际客户请求风格从头撰写而非翻译,由上游维护者的回归测试评判(发布时 withheld)。所有25个修复提交日期在每个评估模型的训练数据截止日期之后,逐任务提供污染论据。我们评估已部署的产品配置(CLI代理+模型+推理工作)——Claude Code与Opus 4.8、Sonnet 5和Haiku 4.5,以及Codex CLI与GPT - 5.5,各独立运行三次,报告任务级置信区间的pass@1、配对比较、美元成本和令牌使用情况。最佳配置解决78.7%的任务;在N = 25时,仅与最弱模型的差距有统计学意义,我们明确指出。审查第五个非竞赛配置(Claude Code + Fable 5,2026年7月2日发布)的完整轨迹时,我们发现产品悄悄替换模型:25个任务中有5个(20%)官方保障回退将常规HTTP协议修复重定向到Opus 4.8——这是已部署产品而非模型是实际测量单元的直接、可重现证据。我们发布任务声明、元数据和完整代理轨迹及差异;评分预言机 withheld,发布时提交SHA - 256清单。

英文摘要

Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue. We introduce RuBench 1.0, a benchmark of 25 tasks mined from recent fix commits in five live open-source repositories (aiohttp, aiogram, Laravel, NestJS, Fastify), each specified natively in Russian -- written from scratch, not translated -- and judged by the upstream maintainer's regression tests, which we withhold from release. All fix commits postdate the training-data cutoffs of every evaluated model. Round 1 evaluates Claude Code with Opus 4.8, Sonnet 5, and Haiku 4.5, and Codex CLI with GPT-5.5 (3 independent runs each; pass@1 with task-level uncertainty); the best configuration resolves 78.7% of tasks. Auditing full trajectories of an hors-concours configuration (Claude Code + Fable 5), we caught the product silently substituting the model on 20% of tasks via an official safeguard fallback -- evidence that the deployed product, not the model, is the unit actually measured. Version 2 adds Round 2: seven further configurations on the same frozen set under a per-configuration freshness gate -- the Russian-market agents SourceCraft CLI (ds, legacy) and Koda CLI (koda-pro), Antigravity with Gemini 3.1 Pro and 3.5 Flash, and Codex CLI with GPT-5.6 Sol and Luna. SourceCraft's flagship resolves 68.1% (N=23), above GPT-5.5 and both Gemini rows. A tool-call contamination re-audit of all 437 Round-2 trajectories finds the Russian and Gemini columns clean (0/293 cells) while flagging systematic oracle-hunting in the GPT-5.6 family (8/69 and 13/75 cells), including one case of mining a prior round's artifacts from the run machine's disk; honest scores are published alongside raw ones. We release statements, metadata, trajectories, and diffs; oracles are withheld with a SHA-256 manifest.

发表机构

  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑