发表机构
Athens University of Economics and Business; Yale University; Rensselaer Polytechnic Institute(雅典经济与商业大学; 耶鲁大学; 伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型推理成本高的问题,提出推理网络框架,证明最优激活策略具有阈值结构,可显著降低成本并满足性能约束。
AI 中文摘要
近年来,大语言模型(LLMs)的进展使其成为自然语言处理(NLP)任务所必需的,而其高昂的推理成本促使人们研究成本-性能权衡。在实践中,多个专家大语言模型被协同用于推理,无论是采用集成模式还是串联模式,但缺乏关于如何最佳利用现有模型的系统性方法。自适应方法可以将简单查询路由到较便宜的LLM,将复杂查询路由到能力更强、成本更高的模型。然而,对于如何最好地利用现有专家模型,目前尚缺乏清晰的理解。我们引入了推理网络,这是一种基于图的框架,其中节点表示不同的LLM,边表示条件模型激活。推理网络设计问题是确定最佳拓扑,即最佳使用模型的方式,以最好地解决成本-性能权衡。我们从一系列LLM专家的基本拓扑开始,每个专家具有不同的成本和不同的专业水平,这通过模型置信度来体现。我们提出了这些模型的最优激活问题,以在满足目标性能约束的前提下最小化预期推理成本。对于这类特殊的推理网络,我们证明了最优激活策略具有阈值结构:首先查询成本最低的LLM,仅当置信度低于定义阈值时才调用更昂贵的LLM。对于判别性任务,最优策略由一组阈值组成,每个类别一个阈值;而对于生成性任务,则由单个阈值组成。我们提供了一种结构化方法来计算阈值,并为两种任务类型提供了实用的置信度估计机制。使用开源LLM进行的实验表明,在满足指定性能预算的同时,实现了显著的成本降低。
英文摘要
Recent advances in large language models (LLMs) have rendered them necessary for Natural Language Processing (NLP) tasks, and their high inference cost motivates the study of cost-performance trade-offs. In practice, several expert LLMs are used in synergy for inference, either in an ensemble mode or in series, yet without a principled approach on how to best use the available models. An adaptive approach can route simple queries to cheaper LLMs and complex ones to more capable, costly models. However, a clear understanding on how to best leverage available expert models is missing. We introduce inference networks, a graph-based framework, where nodes denote different LLMs, and links denote conditional model activations. The inference network design problem is to determine the best topology, namely the best way to use the models that best addresses the cost-performance trade-off. We start from the basic topology of a series of LLM experts, each of which has a different cost and a different level of expertise, which is captured via model confidence. We formulate the problem of optimal activation of these models so as to minimize the expected inference cost subject to a target performance constraint. For this special class of inference networks, we prove that the optimal activation policy has a threshold structure: query the lowest-cost LLM first, and invoke the more expensive LLM only if the confidence falls below a defined threshold. For discriminative tasks, the optimal policy consists of a set of thresholds, one threshold for each class, while for generative tasks, it consists of a single threshold. We provide a structured method to compute the thresholds, and practical confidence estimation mechanisms for both task types. Experiments with open-source LLMs show substantial cost reductions while meeting the specified performance budget.