arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LOCAL:支持智能体大语言模型在设备上进行连续学习

LOCAL: Enabling Learning On-device Contiguously for Agent LLMs

Xinxin Liu, Jiaxin Li, Zibo Wang, Yun Ji, Zhangqi Zhu, Qing Hu, Zhibin Wang, Rong Gu, Sheng Zhong, Chen Tian

arXiv 2608.15241首次发表:更新:

AI 中文总结

LOCAL是首个支持LLM智能体设备端连续学习的单GPU运行时,通过协同组件共享状态实现一致性,在7B级模型、24 GB单GPU上显著降低多项性能指标,保障学习与推理的连续性。

AI 中文摘要

设备端大语言模型(LLM)智能体在本地硬件上与用户反复交互,会产生对自适应调整有价值的私有轨迹,但这些轨迹不应发送给远程训练器。理想情况下,此类智能体应进行连续学习——从每一次交互中自适应调整,且不会暂停或中断面向用户的推理。然而,现有的推理运行时假定权重稳定,现有的强化学习(RL)系统假定资源分离,因此均无法支持这种连续性。我们提出了LOCAL,这是首个支持LLM智能体进行设备端连续学习的单GPU运行时。核心洞见在于,GPU调度、适配器版本管理和KV缓存有效性无法由独立子系统处理:适配器更新会使旧版本的缓存KV张量失效,而缓存保留会影响训练可用的内存。LOCAL使适配器版本、任务优先级和缓存状态对三个协同组件可见——协同调度器、感知版本的KV缓存管理器和多智能体模型运行时,这些组件共享该状态以保持调度、执行和缓存维护的相互一致性。在配备7B级模型的24 GB单GPU上,LOCAL将前台队列等待的第95百分位数(p95)比FIFO降低3.1倍,与不可抢占训练相比将第95百分位数的首token生成时间(TTFT)降低1.55倍,将发布后首次命中的预填充第99百分位数(p99)降低25.6%,跨智能体TTFT的p99降低21.9%,并在严格的KV预算下保持背景学习的推进。

英文摘要

On-device LLM agents interact repeatedly with users on local hardware, producing private traces that are valuable for adaptation but should not be sent to a remote trainer. Ideally, such agents would learn contiguously---adapting from every interaction without pausing or suspending user-facing inference---yet existing inference runtimes assume stable weights and existing RL systems assume separated resources, so neither can support this continuity. We present LOCAL, the first single-GPU runtime that enables contiguous on-device learning for LLM agents. The key insight is that GPU scheduling, adapter version management, and KV-cache validity cannot be handled by independent subsystems: adapter updates invalidate cached KV tensors from older versions, and cache retention affects the memory available for training. LOCAL makes adapter version, task priority, and cache state visible to three cooperating components---a cooperative scheduler, a version-aware KV-cache manager, and a multi-agent model runtime---that share this state to keep scheduling, execution, and cache maintenance mutually consistent. On a single 24 GB GPU with 7B-class models, LOCAL lowers foreground queue-wait p95 by 3.1x over FIFO, lowers p95 time-to-first-token (TTFT) by 1.55x versus non-preemptible training, cuts post-publish first-hit prefill p99 by 25.6% and cross-agent TTFT p99 by 21.9%, and keeps background learning progressing under tight KV budgets.

Comments16 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑