AI 中文总结
研究针对前沿语言模型智能体推理成本高的问题,提出通过自动框架适配让小语言模型智能体以低90%成本匹配语言模型性能,创建相关框架和优化器,经实验验证该方法能扩大小语言模型智能体在日常业务任务中的部署范围。
AI 中文摘要
前沿语言模型智能体正在自动化许多业务任务,但其高推理成本使大规模部署难以持续。小语言模型提供了更便宜的选择,但在替换为为前沿语言模型设计的框架时通常表现不佳。研究表明,对于许多日常业务任务,当小语言模型智能体与元智能体可自动发现的适配框架配对时,能以低90%的成本匹配语言模型性能。关键在于可通过定制指令、工具和编排循环将许多任务难度从模型转移到框架。为此创建了一个框架并构建了框架优化器。在七个面向业务的智能体任务和三个小语言模型系列上的实验表明,优化后的框架显著提高了21个任务-小语言模型对中16个的性能,七对缩小了小语言模型与语言模型的性能差距,最佳小语言模型智能体以4%的成本恢复了89.7%的语言模型性能。分析还表明,适配对于工作流程更重复的任务和具有足够基础能力的小语言模型效果最佳。这些结果表明框架适配可扩大小语言模型智能体在日常业务任务中的实际部署范围。
英文摘要
Frontier LLM agents are automating many business tasks, but their high inference cost makes large-scale deployment unsustainable. Small language models (SLMs) offer a cheaper alternative, yet they typically fall short when swapped into a harness designed for a frontier LLM. We show that for many routine business tasks, SLM agents can match LLM performance at 90% lower cost, when paired with an adapted harness that can be automatically discovered by a meta agent. The key insight is that much of the task difficulty is shared across instances and can be lifted from the model into the harness via tailored instructions, tools, and orchestration loops. To study this systematically, we create a framework that maps agent failure modes to harness adaptation strategies, and build a harness optimizer that automatically discovers effective adaptations from failure trajectories. Across seven business-oriented agentic tasks and three SLM families, we found optimized harnesses significantly improve performance on 16 of 21 task-SLM pairs, with seven pairs closing the SLM-LLM performance gap and the best SLM agent recovering 89.7% of LLM performance at 4% of the cost. Our analysis further shows that adaptation works best for tasks with more repetitive workflows and for SLMs with sufficient base capabilities. Together, these results suggest that harness adaptation can expand the practical deployment range of SLM agents in routine business tasks.