发表机构
LILT, Inc.(LILT公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Terminal-Bench-LILT多语言编码基准,含10种语言300个任务,评估6个前沿模型发现其通过率仅63.1%,凸显多语言编码能力是未被充分探索的维度。
AI 中文摘要
目前针对编码智能体的评估大多仅在英语环境中开展,无法反映真实世界的多语言部署场景。本文提出Terminal-Bench-LILT,这是一个包含10种语言(阿拉伯语、捷克语、德语、西班牙语、印地语、日语、韩语、塞尔维亚语、土耳其语、中文)的300个真实编码任务的基准套件。每个任务都针对非英语软件开发特有的、无直接英语对应项的问题,例如国际化、编码、文本规范化以及文化惯例。所有任务均由母语为对应语言的程序员编写,并通过多阶段质量控制流程验证。对6个前沿模型的评估显示,即使是最强的模型也仅达到63.1%的通过率,还有许多任务未被任何模型解决。不同语言的性能差异显著,且与通用编码基准排名不相关,这表明多语言编码能力是一个独特且未被充分探索的能力维度。示例任务可在指定的URL获取。
英文摘要
Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non-English software development that have no direct English equivalent, e.g., internationalization, encoding, text normalization, and cultural conventions. All tasks are authored by native-speaker programmers and validated through a multi-stage quality control pipeline. Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\% pass rate, with many tasks unsolved by any model. Performance varies substantially by language and does not track general coding benchmark rankings, highlighting that multilingual coding competence is a distinct and underexplored capability axis. Sample tasks are available at https://github.com/lilt/terminal-bench-lilt