发表机构
University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在增强LLM智能体推理能力,提出Agora框架,通过基于拍卖的机制动态分配任务,将推理步骤视为可交易项目让智能体依纠正能力投标,实验表明该框架在多基准测试中表现优且能实现成本-质量权衡。
AI 中文摘要
增强大语言模型(LLM)智能体的推理能力需要有效编排各种专家模型和工具。然而,现有框架通常基于任务与专家模型或工具功能之间的粗粒度匹配来调用应用程序编程接口(API),而忽略了功能相似替代方案之间的性能可变性和成本效率等关键因素。为解决此问题,我们提出了Agora框架,该框架引入了一种激励兼容的拍卖机制,用于将任务动态分配给专家模型和工具。通过将推理步骤视为可交易的项目,Agora使智能体能够根据其纠正后的能力进行投标,确保关键逻辑被路由到最有能力的求解器,而不是最过度自信的求解器。在五个基准测试中的评估表明,在可比候选池下,Agora优于匹配的单模型、路由和级联基线,同时通过单个拍卖参数展现出可控的成本-质量权衡。
英文摘要
Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs, based on coarse-grained matching between tasks and the functions of expert models or tools, while overlooking critical factors such as performance variability and cost efficiency among functionally similar alternatives. To address this, we propose Agora, a framework that uses a confidence-calibrated auction to dynamically allocate tasks to expert models and tools. By treating reasoning steps as tradeable items, Agora bases allocation on calibrated competence rather than raw confidence. Across five main benchmarks, Agora improves or remains competitive with single-model, routing, and cascade baselines under matched candidate pools.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026. 13 pages, 5 figures