内存主权推理:超越完全驻留的输出精确执行
Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency
AI总结:
该研究针对存储支持推理的夸大问题,提出可证伪证书,通过Qwen3-Next等实验验证,实现超越完全驻留的输出精确执行,异步组件耗时仅为阻塞直接路径的32.3%。
AI中文摘要:
存储支持的推理容易被夸大:进程RSS不包含计费页缓存,进程级设备读数不包含板级使用情况,且成功生成不代表正确的异步复用。我们提出一种可证伪的证书,用于分离表示、语义需求、调度器请求和流量,同时命名资源权限、精确性范围和已测试的复用转换。在一个Qwen3-Next实例中,标准路由器在32K预填充期间选择全部48个×512的托管层专家对象,它们无重复、无重叠的规范并集给出43.59375 GiB的语义需求下界,超过声明的34 GiB完全驻留包络和11 GiB主机硬件加24 GiB设备的包络。LRU64执行保持在其主机硬件/GPU审计的契约内。针对一个预先指定的零缓存预言机,全部64个标记ID、64个完整的151936浮点数logit行的每个字节、3408个路由事件、响应字节以及记录的消费者和目标身份均为精确,排除循环状态和上游运行时相等性。在匹配的源任务中,缓冲路径完全完成,但达到11 GiB主机上限并记录33481个事件。阻塞单窗口直接路径和完整八窗口异步组件具有正余量且无极限事件。在六个平衡对中,完整异步组件在每输出相同物理源字节的情况下,占用阻塞直接路径墙钟时间的32.3%,该比较同时改变队列深度、重叠和生命周期实现。此外,14个预先指定的控制/故障单元在后续Qwen3-Next和Gemma 4二进制文件中通过,仅对指定转换支持故障关闭行为。主要实验使用固定模型、工作负载、运行时、设备和64输出范围。
英文摘要:
Storage-backed inference is easy to overclaim: process RSS excludes charged page cache, process-local device readings exclude board-wide use, and successful generation does not establish correct asynchronous reuse. We present a falsifiable certificate separating representation, semantic demand, scheduler requests, and traffic while naming resource authorities, exactness horizons, and tested reuse transitions. At one Qwen3-Next identity, the stock router selects all 48 x 512 managed layer-expert objects during 32K prefill. Their duplicate-free, overlap-free canonical union gives a 43.59375 GiB semantic-demand lower bound, exceeding the declared 34 GiB full-residency envelope and the 11 GiB host-hard plus physical 24 GiB-device envelope. LRU64 execution stays within its host-hard/GPU-audited contract. Against one prespecified zero-cache oracle, all 64 token IDs, every byte of 64 complete 151,936-float logit rows, 3,408 route events, response bytes, and recorded consumer and destination identities are exact. Recurrent-state and upstream-runtime equality are excluded. In a matched source campaign, the buffered path completes exactly but reaches the 11 GiB host ceiling and records 33,481 memory.max events. The blocking one-window direct path and complete eight-window asynchronous component are exact with positive margin and zero limit events. Across six counterbalanced pairs, the complete asynchronous component takes 32.3% of the blocking direct path's wall time at identical physical source bytes per output; the comparison jointly changes queue depth, overlap, and lifecycle implementation. Separately, fourteen prespecified control/fault cells pass across later Qwen3-Next and Gemma 4 binaries, supporting fail-closed behavior only for the named transitions. The principal experiment is one fixed model, workload, runtime, device, and 64-output horizon.