AI 中文总结
该研究针对LLM智能体控制,提出就绪队列边界的形式化方法,通过实验验证设备驻留路由决策路径可提升效率,为GPU智能体控制确立了两个可测量门限。
AI 中文摘要
LLM智能体服务会在模型调用与工具调用之间反复执行小型确定性转换:路由结果、更新状态、发出下一个效果。我们探究该控制路径何时能暴露足够的并发工作以适配GPU执行,以及当GPU计算的路由决策保留在设备上时会产生何种变化。我们使用固定分区份额F、精确离线份额P*、局部上界U和在线达成份额A来形式化就绪队列边界。在零服务时间、无限容量及相等相对启动截止期限的条件下,专用动态规划可精确计算P*。在一个固定的851会话公开轨迹面板的平稳泊松重放中,当目标活跃会话数为100000、K=256且启动截止期限为50ms时,得到F=30.19%、P*=43.00%、U=45.85%。精确打包可恢复固定窗口边界下损失机会的81.83%。源于结果的路由密钥是一种条件代理,而非可执行身份的证明。另一项机制研究将GPU计算的二进制决策保留在设备上,而非向主机返回4字节并重新调度。在四个指定的GPU部署中,设备驻留路径在全部36种配置下均更快;同一部署内的行中位数比率介于1.19倍至2.39倍之间。在两种可容许机制下,全部14557440次测试的批量调用均与单独实现的主机预言机匹配。不删除任何主机决策的固定嵌套设备图在五个部署的全部60种配置下均更慢。综上,本研究为GPU智能体控制确立了两个可测量的门限:截止期限可行的队列供给与观测放置。需加入有限在线运行时以测量A、CPU位移及服务级效益。
英文摘要
We introduce a mechanism that improves LLM-agent execution: keeping four-byte, GPU-computed control-path routing decisions on-device, avoiding a round-trip to host memory for redispatch. It is faster than host-dispatch in all 36 placements x mechanisms settings in our benchmark across four named GPUs with row-median speedup of 1.19x-2.39x, and outputs are correct, matching an independent host implementation, for all 14,557,440 calls in all configurations, in both tested mechanisms. Agent control paths can be deterministic. Their transitions between LLM calls and external tool calls can be GPU-accelerated by batching across agents and time. We define a cohort of transitions with launch times in a window, and study four shares of GPU-active time: a fixed-partition baseline F, an optimal P*, an upper bound U, and an online result A. Under assumptions of uniform relative launch deadlines, infinite capacity, and zero service times, we compute P* with a custom dynamic program and validate our simulator against it. We replay an anonymized 851-session trace, shown in full as a panel, stationary-ized with a Poisson process. With 100,000 target active agents and a launch deadline of 50 ms, where the headline results have K=256, we find F=30.19%, P*=43.00%, U=45.85%; our method for computing P* reveals that 81.83% of this opportunity is recoverable from cohort packing at fixed window ends. Our method for U relies on a novel notion of route-key, used only as a proxy (via conditional statistics) and does not purport to establish route-equivalence. We find a nested graph that retains some host-side decisions slower in all 60 configurations (five placements).
Comments14 pages, 4 figures. Includes formal proofs, trace provenance, and a reproducibility appendix. Code and artifacts: https://github.com/josefchen/ready-cohorts ; processed evidence: https://huggingface.co/datasets/josefchen/ready-