AI 中文总结
针对LLM智能体循环中KV缓存驻留管理问题,提出UNISON近内存调度器,通过SPEAR和TIDE联合策略提升命中率并降低延迟,在多个基准上验证了有效性。
AI 中文摘要
大型语言模型日益被组合成智能体循环,这些循环会规划、调用工具,并在每次操作后恢复同一任务。这些循环对共享内存层次结构的压力比传统多轮对话更大,因为它们在工具等待期间持有不断增长的键值(KV)前缀,并将许多会话放置在一个SRAM/HBM池中,因此驱逐和层次化放置成为一个会话级别的效率问题,与计算模式优化正交。现有的基于最近使用、超时或身份的代理机制忽略了循环的机制信息,因此将活跃等待视为冷数据、可丢弃单元。我们提出了统一原生轮间会话编排联结(UNISON),一种事件驱动的近内存调度器,其中用于智能体返回间隙的生存惩罚驱逐(SPEAR)和空闲窗口DMA事件中的分层(TIDE)共享一个实时排名。SPEAR根据间隙平均值和轮次索引危险度选择谁离开,而TIDE将观察到的等待作为DMA预算用于谁留在快速层。在三个模型家族、总计1,415个会话和33,596轮的编码和通用任务基准上,联合策略是每条轨迹上最佳的非预言机条目,将命中率提高0.3%至23.1%,将平均内存访问时间(AMAT)降低22%至51%,并在长时程轨迹上将首令牌时间(TTFT)降低58%至89%。结构必要性分析表明,统一近内存设计无法分解为独立IP,也无法在软件中实现而不重新引入已记录的故障模式。28纳米CMOS调度核心占用0.169平方毫米,功耗13.6毫瓦,频率150兆赫,相对于其管理的KV层次结构开销可忽略不计,其浮点排名与Kendall tau超过0.998的参考值一致。
英文摘要
Large language models are increasingly composed into agent loops that plan, call tools, and resume the same task after each action. These loops press a shared memory hierarchy harder than conventional multi-turn chat, because they hold a growing key-value (KV) prefix across tool waits and place many sessions on one SRAM/HBM pool, so that eviction and hierarchical placement become a session-level efficiency problem orthogonal to compute-mode optimization. Existing proxies based on recency, timeout, or identity miss the mechanism information of the loop and therefore treat a live wait as a cold, discardable unit. We present Unified Native Inter-turn Session Orchestration Nexus (UNISON), an event-driven near-memory scheduler in which Survival-Penalty Eviction for Agent Return-gap (SPEAR) and Tiering in Idle-window DMA Events (TIDE) share one live ranking. SPEAR selects who leaves from a gap average and a turn-indexed hazard, while TIDE spends the observed wait as a DMA budget for who sits in the fast tier. On coding and general-mission benchmarks with three model families, totaling 1,415 sessions and 33,596 turns, the joint policy is the best non-oracle entry on every trace, raising hit rate by 0.3% to 23.1%, reducing AMAT by 22% to 51%, and lowering TTFT by 58% to 89% on long-horizon traces. A structural necessity analysis shows that the unified near-memory design cannot be decomposed into independent IPs or realized in software without re-introducing documented failure modes. The 28-nm CMOS scheduling core occupies 0.169 mm^2 at 13.6 mW and 150 MHz, a negligible overhead relative to the KV hierarchy it manages, reproducing the floating-point ranking at Kendall tau exceeding 0.998.