发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体的工具链配置问题,提出映射引导的升级算法,在液体冷却任务中提升了准确率并减少token使用,在电网任务中提供低成本替代方案,揭示工具链配置遵循领域依赖的帕累托前沿。
AI 中文摘要
LLM智能体已被广泛应用于关键任务基础设施(MCI)的运行中,这些智能体通常依赖工具链(harness)来确定其可访问的信息、可使用的工具以及可执行的动作。现有系统往往为所有任务提供相同的全面工具链,这可能并非必要,还会造成资源浪费。在本文中,我们专注于最优工具链配置的识别,将其视为每个任务需求与工具链提供能力之间的资源匹配问题。为衡量这种匹配度,我们基于底层系统的数学表示对MCI任务进行分类,并按工具链提供的信息数量和类型对其配置进行排名。随后,我们从两个来源构建任务到工具链的映射:挖掘研究文献和测量受控的智能体执行。利用测得的映射,我们提出一种新的工具链配置算法:映射引导的升级(map-guided escalation)。该算法从特定任务的工具链开始,仅在自我检查失败后才扩展至全面配置。我们在两个代表性MCI任务中评估了我们的方法:在液体冷却任务中,它将智能体的准确率从全面配置下的0.652提升至0.715,且在准确率与Reflexion相当的情况下,使用的token减少了48%;在电网任务中,全面配置仍为准确率最优选择,而基于映射的配置则提供了成本更低的替代方案。这些发现表明,工具链配置遵循依赖于领域的准确率-成本帕累托前沿,而非存在通用最优解。
英文摘要
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the identification of optimal harness configurations, and view it as a resource-matching problem between what each task requires and what the harness provides. To measure this match, we classify MCI tasks based on the mathematical representation of the underlying system and rank harness configurations by the amount and type of information they provide. We then construct task-to-harness mappings from two sources: mining research literature and measuring controlled agent execution. Leveraging the measured mapping, we propose a new harness provisioning algorithm: map-guided escalation. It begins with a task-specific harness and expands to full provision only after a failed self-check. We evaluate our method in two representative MCI tasks: in liquid cooling, it improves the agent accuracy from 0.652 under full provision to 0.715 and achieves accuracy comparable to Reflexion with 48% fewer tokens; In power grids, full provision remains accuracy-optimal, while map-based provisioning offers lower-cost alternatives. These findings show that harness provisioning follows a domain-dependent accuracy-cost Pareto frontier rather than a universal optimum.