arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Opti-Agent-Bench:在实际业务问题上对端到端优化研发智能体进行基准测试

Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

Yongchang Fu, Xinjie Huang, Chengjun Dai, Chengzhe Feng, Junshao Zhang, Hong Zhu

arXiv 2607.10768首次发表:更新:

发表机构

Ding Talk, Alibaba Group; Zhejiang University(钉钉,阿里巴巴集团; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对大语言模型解决优化问题的现有不足,引入Opti-Agent-Bench基准测试,通过业务语义真实性、模块化评估、ORAC双层有效性框架,在多工业规模任务中揭示当前模型关键失败模式。

AI 中文摘要

基于大语言模型的智能体越来越多地用于解决优化问题,但现有基准测试是在预结构化数学公式上评估它们,绕过了最关键挑战:将复杂业务需求转化为正确模型并高效求解。我们引入Opti-Agent-Bench,这是一个端到端基准测试,评估大语言模型在完整优化研发流程中的表现,从理解业务语言描述到数学建模、算法选择、代码实现,再到生成解决方案报告。其设计基于三个支柱:具有反模板陷阱的业务语义真实性、跨模块一致性检查的模块化评估、确保任务质量和评分完整性的ORAC双层有效性框架。在多个工业规模任务中,揭示了当前模型的关键失败模式。

英文摘要

LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently. We introduce Opti-Agent-Bench, an end-to-end benchmark that evaluates Large Language Models (LLMs) across the complete optimization R&D pipeline, from understanding business-language descriptions through mathematical modeling, algorithm selection, and code implementation, to solution report generation. Our design rests on three pillars: (1) businesssemantic authenticity with anti-template traps that defeat pattern matching; (2) modular evaluation with cross-module consistency checking across Problem Understanding, Formal Modeling, Implementation, and Reporting; and (3) the ORAC bi-level validity framework that simultaneously ensures task quality and scoring integrity. Across several industrialscale tasks spanning integer programming, robust optimization, stochastic programming, and non-convex optimization, we expose critical failure modes of current models, including constraint omission, model-code inconsistency, and report-implementation divergence, that remain invisible under conventional single-metric evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑