arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33141cs.AIcs.SE

设备端智能体操作缓存——以分类器为中心的从自然语言到动作生成

On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation

Moghis Fereidouni, Anthony Arnold, Sumit Gulwani, Mark Marron, A. B. Siddique

AI总结:

本文提出设备端操作缓存,将NL-to-Action从生成式转为分类式,在设备端处理常见动作,降低延迟、增强隐私、降低成本,在Excel公式生成任务中推理成本降低56%,缓存命中延迟降低5倍。

AI中文摘要:

智能体人工智能正日益嵌入软件应用中,以提供面向功能和特性的自然语言界面。在大多数情况下,这些智能体由企业级(超过1000亿参数)或前沿级大语言模型驱动,这些模型需要大量的计算资源来运行,并依赖云端托管的推理来处理将自然语言输入转换为可操作的软件操作的任务。这种对云端托管推理的依赖,在LLM推理时间之上引入了大量的网络延迟,引发了数据隐私担忧,并且,考虑到运行这些模型的成本,可能会迅速增加与支持智能体特性相关的开销。本文提出了一种新颖的方法,通过设备端操作缓存,将自然语言到动作(NL-to-Action)问题从生成式表述转变为以分类器为中心的表述。这些缓存使得智能体系统能够完全在设备端处理频繁出现的动作类别——从而减少延迟、增强隐私并降低运营成本。我们表明,对于经典的NL-to-公式任务,即根据用户请求生成Excel公式,与仅依赖云端的模型路由推理相比,该方法将总推理成本降低了56%,并且在缓存命中时,将响应延迟降低了5倍。

英文摘要:

Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resources run and depend on cloud hosted inference to handle the task of transforming natural language inputs into actionable software operations. This reliance on cloud-hosted inference introduces substantial network latency on top of LLM inference times, creates data privacy concerns, and, given the costs of running these models, can rapidly escalate expenses associated with supporting agentic features. This paper introduces a novel means of converting the NL-to-Action problem from a generative one into a classification-centric formulation via on-device operation caches. These caches allow an agentic system to handle frequently occurring classes of actions completely on-device -- reducing latency, enhancing privacy, and lowering operational costs. We show that for a classic NL-to-Formula task, generating Excel Formula in response to user requests, this approach reduces total inference cost by 56% when compared to cloud-only model-routing based inference and, on cache hits, reduces the latency to response latency by 5x.

↑