AKTS:面向语言模型智能体的亚微秒级内核策略切换
AKTS: Sub-Microsecond Kernel Policy Switching for Language-Model Agents
浏览论文内容
中文总结 AI 辅助
AKTS通过预验证策略库和索引切换,实现亚微秒级内核调度策略切换,兼顾突发响应与后台吞吐,在vLLM中捕获97%批量工作。
中文摘要 AI 辅助
基于GPU的LLM服务器通常在同一CPU上复用交互式请求与后台批量工作。在请求突发期间,调度器应保护首令牌时间;在突发之间,应让后台工作取得进展。固定的内核策略会牺牲其中一个目标,因此智能体操作系统控制需要一种随工作负载变化而切换调度器行为的方法。难点不在于判断切换是否有用,而在于安全且足够快地应用于内核。调度器事件每1-10微秒发生一次,任何在该处运行的代码都必须满足eBPF验证器。标量旋钮速度快,但仅暴露有限的策略行为,而生成新的eBPF策略代码具有表现力,但将编译、验证、加载以及可能的验证器拒绝置于运行时路径上。我们提出AKTS,它在加载时一次性验证策略库,并将智能体的运行时动作简化为向内核内预验证策略数组写入整数索引。内核内尾调用解析该索引。由于智能体发出索引而非代码,验证器失败不再是运行时结果。在Linux 6.14上,AKTS在920纳秒(p50)内应用策略切换,与标量写入匹配的同时切换整个策略;使无效索引在附加调度器上的60,217次调用中保持惰性;并在vLLM工作负载中切换策略,捕获吞吐量策略97%的批量工作,同时匹配延迟策略的突发响应。
英文摘要
GPU-backed LLM servers multiplex interactive requests with background batch work on the same CPUs; a fixed kernel policy serves one objective and loses the other, so agentic OS control needs to switch scheduler behavior as the workload changes. The hard part is applying the switch safely and fast enough for the kernel: scheduler events occur every 1-10 $μ$s, and any code running there must satisfy the eBPF verifier. Scalar knobs are fast but limited, while generating eBPF policy code puts compilation, verification and possible rejection on the runtime path. We present AKTS, which verifies a policy library once, at load time, and reduces the agent's runtime action to writing an integer index into an in-kernel array of preverified policies, resolved by a tail call. Because the agent emits an index rather than code, verifier failure is not a runtime outcome. On Linux 6.14, AKTS applies a policy switch in 920 ns (p50); makes an invalid index inert across 60,217 invocations on an attached scheduler; and switches policies in a vLLM workload to capture 97% of a throughput policy's batch work while matching a latency policy's burst response.
发表机构
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。