发表机构
Harvard; MIT(哈佛大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM部署挑战,提出OrchSLM框架统一非交互式SLM编排方法,揭示任务结构、模型池与共识等参数如何影响编排行为。
AI 中文摘要
尽管大型语言模型(LLMs)已展现出卓越的能力,但它们对云规模基础设施的依赖给其在智能体管道中的部署带来了根本性挑战,包括延迟、隐私、连接性和高昂的计算成本。小语言模型(SLMs)提供了一种引人注目的替代方案:近期研究表明,在智能体工作负载中,许多重复且范围狭窄的子任务可能更适合由专门的SLMs而非单一的大型LLM来处理。然而,SLMs有限的容量和上下文窗口可能限制其进行长时程推理以及诸如迭代验证和辩论等交互密集型编排策略的能力。这促使我们提出一种互补的、非交互式的范式,在该范式中,异构SLMs独立生成候选解决方案,而一个路由器在不进行进一步模型交互的情况下编排其缓存样本。为了进一步理解这种编排的机制,我们引入了OrchSLM,一个路由框架,它统一了现有的非交互式编排方法,并将其底层设计选择暴露为可控参数。利用OrchSLM作为系统性探针,我们揭示了编排行为如何从包括任务结构、模型池组成和多智能体共识在内的多种旋钮中涌现。
英文摘要
Although large language models (LLMs) have demonstrated remarkable capabilities, their reliance on cloud-scale infrastructure poses fundamental challenges for deployment in agentic pipelines, including latency, privacy, connectivity, and substantial computational cost. Small language models (SLMs) offer a compelling alternative: recent studies suggest that many repetitive and narrowly scoped subtasks in agentic workloads may be better served by specialized SLMs than by monolithic LLMs. However, the limited capacity and context windows of SLMs can constrain long-horizon reasoning and interaction-heavy orchestration strategies such as iterative verification and debate. This motivates a complementary, non-interactive paradigm in which heterogeneous SLMs independently generate candidate solutions and a router orchestrates their cached samples without further model interaction. To further understand the mechanisms of such orchestration, we introduce OrchSLM, a routing framework that unifies existing non-interactive orchestration methods and exposes their underlying design choices as controllable parameters. Using OrchSLM as a systematic probe, we reveal how orchestration behavior emerges from diverse knobs, including the task structure, model-pool composition, and multi-agent consensus.