MiniCache:用于高效大语言模型推理的具有小模型接口的可重复使用程序缓存
CacheSpec: Finding the Sweet Spot for Small Models in Large Language Models
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型推理成本高的问题,提出MiniCache框架,将PoT程序转换为参数化缓存对象,通过重用小模型进行语义变量提取和推测性起草,减少目标大语言模型调用,实验证明其能提升推理性能,平衡延迟、缓存重用和准确性。
中文摘要 AI 辅助
大语言模型(LLMs)越来越多地用于程序辅助推理、智能决策和结构化任务执行,但这些应用通常会产生高昂的推理成本。我们提出了MiniCache,一个可重复使用的程序缓存框架,它将思维程序(PoT)转换为参数化缓存对象,实现跨结构相似请求的可重复计算。MiniCache在缓存命中请求时重用相同的小模型进行语义变量提取,并在目标大语言模型生成时进行推测性起草,减少昂贵的目标大语言模型调用,同时保持任务质量。在购物风格请求数据集、WebShop、Formula和CodeTAT-QA上的实验表明,MiniCache在推理延迟、缓存重用和准确性之间取得了更好的平衡,在并行服务下延迟降低了3.1倍,吞吐量提高了2.8倍。这些结果表明,小模型最有效的作用不是替代大模型,而是作为轻量级接口模型,实现可靠且高效的可重复使用程序缓存。
英文摘要
Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these settings often incur substantial inference cost. Many such requests share similar computational structures while differing in variables, constraints, or contexts, creating opportunities for program-level caching. Since program caches need to reapply reusable computation logic to new requests, their key steps often involve lightweight and structured operations such as variable extraction, program binding, and generation acceleration, which are well suited for small models. We propose CacheSpec, an inference optimization framework centered on reusable program caches. The framework converts Program-of-Thoughts (PoT)-style programs from one-time reasoning artifacts into reusable cache objects, and reuses the same small model for two roles: semantic variable extraction on the cache-hit path and speculative drafting during target-LLM generation. Experiments on shopping-style request datasets, WebShop, Formula, and CodeTAT-QA show that CacheSpec reduces inference latency and improves effective cache reuse while preserving comparable or better task quality than existing caching and generation baselines, achieving up to about 3.1$\times$ latency speedup; in parallel serving experiments, it improves throughput by about 2.8$\times$ over PoT-style methods. These results suggest that the sweet spot for small models in large-model inference systems lies not in solving complex tasks independently, but in performing lightweight, structured, and verifiable auxiliary operations.