arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlgoWorlds:用于算法世界中全局优化的工具使用基准

AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds

Zixiang Xu, Jiaan Wang, Fandong Meng

arXiv 2608.29397首次发表:更新:

发表机构

Weixin AI, Tencent; University of Southern California(腾讯微信AI; 南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出AlgoWorlds基准,评估大语言模型在算法世界中利用工具将信息转化为全局最优决策的能力,实验显示领先模型仅38.61%案例能达精确最优。

AI 中文摘要

工具使用基准通常评估智能体是否使用合适的工具和有效参数完成工作流程。然而,仅可行性在路线规划、车队调度等现实决策场景中是不够的,个体选择会通过共享约束和成本产生相互作用,因此可行的解决方案仍可能严重次优。这提出了一个更具挑战性的问题:智能体能否将通过工具收集的信息转化为全局最优决策?我们推出AlgoWorlds,这一基准将形式化指定的组合优化问题转化为具有可验证全局最优解的部分可观测决策环境。每个环境包含仅能通过特定任务信息工具观测的隐藏实例,之后智能体需做出一个结构化决策,该决策会针对可行性和最优性进行评估。AlgoWorlds包含240个环境,覆盖10个组合优化家族和4个工作负载级别;特定家族的确定性程序生成实例,精确算法验证其最优性并确定工作负载级别,两种结构不同的工具界面呈现每个底层实例。我们评估了7种领先的大语言模型(LLMs),包括Claude Opus 4.8和GPT-5.6 Sol。实现全局最优性仍极具挑战性:尽管领先模型在大多数情况下能生成可行决策,但表现最佳的模型仅在38.61%的案例中达到精确最优。即使智能体收集到足够信息以重构隐藏实例,多数失败仍以可行但次优的决策告终。因此,该挑战不仅限于信息获取,还延伸至信息整合、全局约束推理和决策验证。项目主页和代码可在指定URL获取。

英文摘要

Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is insufficient in real-world decision settings such as route planning and fleet dispatch. Individual choices interact through shared constraints and costs, so a feasible solution may still be substantially suboptimal. This raises a harder question: can an agent turn information gathered through tools into a globally optimal decision? We introduce AlgoWorlds, a benchmark that transforms formally specified combinatorial optimization problems into partially observed decision environments with verifiable global optima. Each environment contains a hidden instance observed only through task-specific information tools, after which the agent commits to one structured decision evaluated for feasibility and optimality. AlgoWorlds contains 240 environments covering ten combinatorial optimization families and four workload levels. Family-specific deterministic programs generate the instances, exact algorithms certify their optima and determine workload levels, and two structurally different tool interfaces present each underlying instance. We evaluate seven leading LLMs, including Claude Opus 4.8 and GPT-5.6 Sol. Achieving global optimality remains highly challenging: although leading models produce feasible decisions in most cases, the best-performing model reaches exact optimality in only 38.61% of cases. Even when agents collect sufficient information to reconstruct the hidden instance, most failures end in feasible but suboptimal decisions. The challenge therefore extends beyond information acquisition to information integration, global constraint reasoning, and decision verification. The project homepage is available at https://xzx34.github.io/AlgoWorlds/, and the code is available at https://github.com/xzx34/AlgoWorlds.

CommentsHomepage: https://xzx34.github.io/AlgoWorlds/ Code: https://github.com/xzx34/AlgoWorlds

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑