arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Terminal-Bench-LILT:基于语言、区域与文化的多语言智能体编码基准

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

Yunsu Kim, Kaden Uhlig, Ashwin Purohit, Milind Agarwal, Patrick Simianer, Anil Arslan, Kiarash Mokhtari, Thomas Zenkel, Johannes Mosig, Gabriel Bretschner, Shamik Bose, Joern Wuebker, John DeNero

arXiv 2608.28641首次发表:更新:

发表机构

LILT, Inc.(LILT公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Terminal-Bench-LILT多语言编码基准,含10种语言300个任务,评估6个前沿模型发现其通过率仅63.1%,凸显多语言编码能力是未被充分探索的维度。

AI 中文摘要

目前针对编码智能体的评估大多仅在英语环境中开展,无法反映真实世界的多语言部署场景。本文提出Terminal-Bench-LILT,这是一个包含10种语言(阿拉伯语、捷克语、德语、西班牙语、印地语、日语、韩语、塞尔维亚语、土耳其语、中文)的300个真实编码任务的基准套件。每个任务都针对非英语软件开发特有的、无直接英语对应项的问题,例如国际化、编码、文本规范化以及文化惯例。所有任务均由母语为对应语言的程序员编写,并通过多阶段质量控制流程验证。对6个前沿模型的评估显示,即使是最强的模型也仅达到63.1%的通过率,还有许多任务未被任何模型解决。不同语言的性能差异显著,且与通用编码基准排名不相关,这表明多语言编码能力是一个独特且未被充分探索的能力维度。示例任务可在指定的URL获取。

英文摘要

Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish, Hindi, Japanese, Korean, Serbian, Turkish, and Chinese. Each task targets issues specific to non-English software development that have no direct English equivalent, e.g., internationalization, encoding, text normalization, and cultural conventions. All tasks are authored by native-speaker programmers and validated through a multi-stage quality control pipeline. Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\% pass rate, with many tasks unsolved by any model. Performance varies substantially by language and does not track general coding benchmark rankings, highlighting that multilingual coding competence is a distinct and underexplored capability axis. Sample tasks are available at https://github.com/lilt/terminal-bench-lilt

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑