arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16777cs.CLcs.AI

JOR-Bench:用于大语言模型的日语运筹学基准测试

JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models

Yuu Jinnai

首次发表
浏览论文内容

中文总结 AI 辅助

JOR-Bench是用于评估大语言模型解决运筹学问题能力的日语基准测试集,涵盖多种规划问题。通过评估七个LLMs的英文和日语版本,发现多语言模型的OR制定能力语言中立,但存在跨语言差异。

中文摘要 AI 辅助

我们展示了JOR-Bench,这是一个包含五个日语基准测试的集合,用于评估大语言模型(LLMs)制定和解决运筹学(OR)问题的能力。每个基准测试都是现有英文基准测试的日语翻译,涵盖线性规划、混合整数规划、非线性规划和组合优化等1319个问题。JOR-Bench是一个独立于求解器的基准测试,可与任何求解器或编程语言一起使用。我们在原始英文和新日语版本上评估了七个LLMs,并比较了跨语言的性能。结果表明,对于强大的多语言模型,OR制定能力在很大程度上是语言中立的,但错误分析揭示了微妙的跨语言差异。

英文摘要

We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formulate and solve operations research (OR) problems. Each benchmark is a Japanese translation of an existing English benchmark: IndustryOR, MAMO Complex LP, NL4OPT, OptiBench, and OptMATH, covering 1,319 problems spanning linear programming, mixed-integer programming, non-linear programming, and combinatorial optimization. JOR-Bench is a solver-independent benchmark that can be used with any solver or programming language, and consists of pairs of Japanese problem statements and expected numerical answers. We evaluate seven LLMs, including multilingual general-purpose models and Japanese-specialized models, on both the original English and the new Japanese versions, and compare performance across languages. For the main evaluation, we standardize execution with the Python interface to OR-Tools to make model outputs comparable and reproducible with open-source software. Our results show that OR formulation ability is largely language-neutral for strong multilingual models; the overall average accuracy difference between English and Japanese is only $-0.3$ pp. Yet error analysis reveals subtle cross-lingual differences, including a pragmatic disambiguation failure in some domains that causes models to output decision-variable values instead of the objective value when the prompt is in Japanese.

发表机构

  • CyberAgent

机构由 AI 辅助整理,请以论文原文为准。

↑