arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20397cs.AI

Nexus:面向统一内存上智能体大语言模型的深度自适应KV缓存拼接与检索解耦工具路由

Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory

Mustafa Arslan

首次发表
浏览论文内容

中文总结 AI 辅助

Nexus针对MCP智能体LLM工具模式预填充主导TTFT的问题,通过检索解耦路由与预填充成本,结合深度自适应KV缓存拼接技术,实现了工具选择与参数生成的高效处理,提升了首token生成速度与上下文利用率。

中文摘要 AI 辅助

基于模型上下文协议(MCP)的智能体大语言模型(LLM)每轮都会重新编码冗长的工具模式,因此随着工具注册量增长,序列长度二次方级别的预填充会主导首token生成时间(TTFT)。Nexus的核心手段是将路由与模式预填充成本解耦:INT8语义旁加载缓冲区(SLB)通过校准后的交叉编码器边距门控,以检索方式选择工具,且参数生成基于压缩文本签名(中位数19个token)而非拼接的键值(KV)缓存,该路径与深度无关;当工具注册量扩展至250个时,路由准确率保持在近89%,而拼接所有模式的基准方法会完全溢出上下文窗口,且其首参数token生成速度比全模式重新预填充快1.66倍,同时主上下文token节省约80%。作为次要的有限手段,Nexus会将编译后的模式KV块直接移植到活跃上下文,但该操作本质上受旋转位置嵌入(RoPE)相位漂移限制:锚定拼接的输出完全准确,而偏离锚定的放置会破坏注意力,因此在超过阈值P=256时,Nexus会通过深度自适应后缀重解码修复拼接处,必要时升级为全量重新预填充;这种永不退化特性是输出保真度的保证(Top-1一致性,KL散度D_KL≈0),而非延迟保证,延迟可降至0.98倍后收敛至相等,同时在中等深度下TTFT加速1.1-1.7倍,在深度上下文下收敛至相等。两项负面结果限定了该设计:偏离锚定的RoPE保真度边界,以及无参考漂移门无法预测漂移(斯皮尔曼相关系数rho=0.193)。所有测量均来自苹果硅统一内存上的一个模型组合(Qwen2.5-14B-Instruct Q4_K_M);定性边界具有通用性,而定量范围与该组合相关。

英文摘要

Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to decouple routing from the schema-prefill cost: an INT8 semantic lookaside buffer (SLB) with a calibrated cross-encoder margin gate selects tools by retrieval, and arguments are generated over a compressed textual signature (median 19 tokens) rather than over spliced key/value (KV) cache. This path is depth-independent: routing accuracy stays near 89% as the registry scales to 250 tools - where a concatenate-all-schemas baseline overflows the context window entirely - and it reaches a first-argument token 1.66x sooner than a full-schema re-prefill at a ~80% main-context token saving. As a secondary, bounded lever we transplant a compiled schema KV block directly into the live context. This is fundamentally limited by rotary position embedding (RoPE) phase drift: an anchored splice is output-exact, but off-anchor placement corrupts attention, so beyond a threshold P=256 Nexus repairs the seam with a depth-adaptive suffix redecode that escalates to a full re-prefill. The resulting never-regress property is a guarantee on output fidelity (top-1 agreement, D_KL approx. 0) - not on latency, which can dip to 0.98x before converging to parity - alongside a 1.1-1.7x TTFT speedup at moderate depth that narrows to parity at deep context. Two negative results bound the design: the off-anchor RoPE fidelity boundary, and the failure of a reference-free drift gate to predict drift (Spearman rho = 0.193). All measurements are from one model tuple (Qwen2.5-14B-Instruct Q4_K_M) on Apple-silicon unified memory; the qualitative boundaries generalize, while the quantitative envelope is tuple-specific.

↑