AI 中文总结
JustFit 通过 KVExec、PhaseSwap 和 StateTrans 机制实现即时状态管理,在 24 GiB 笔记本上支持 200K token 服务,容量提升 6.93 倍,并保持高准确率。
AI 中文摘要
能力强大的开放权重模型使得本地编码和推理变得可行,但其上下文和执行状态给笔记本电脑内存带来压力。我们提出了 JustFit,一个基于 MLX 的推理运行时,它结合了用于压缩 KV 执行的 KVExec、用于组件驻留的 PhaseSwap 以及用于状态保持服务切换的 StateTrans。这些机制融合了重建,并独立于模型权重量化,协调即时物化和释放。在配备 Qwen3.8-27B MXFP4 的 24 GiB M4 Pro MacBook 上的全执行容量测试中,三次独立运行完成了 196,608 个输入和 16,384 个输出 token,将单请求上下文从 mlx-vlm 基线的 30,720 个位置增加到 212,992 个(6.93 倍);另一次双请求运行合计保留了 229,376 个位置。在单独的性能测试中,一个 32K 输入、64 输出的探针达到 19.11 token/s,而重复的 32K+6K 工作负载的中位峰值进程占用为 16,374 MiB。集成运行时正确回答了 AIME 2026 的 30 个问题中的 29 个,展示了紧凑状态和生命周期感知执行如何扩展本地服务容量,同时支持扩展的生成推理。
英文摘要
Local agents need memory for model execution and working history. We present JustFit, an MLX runtime that coordinates their overlapping allocations: KVExec executes and checkpoints four-bit KV with bounded workspace, PhaseSwap loads phase-dependent components, and StateTrans preserves history across execution modes. With Qwen3.8-27B MXFP4 on a 24 GiB M4 Pro MacBook, JustFit reaches 327,680 retained positions across two requests, 10.67 times the evaluated baseline's 30,720-position single-request record. One request completes the full 262,144-position native window at median 5.986 tokens/s. Each shape completes 16,384 outputs per request in three fresh processes: B1 uses one cold build and two prefix extensions; B2 uses three ordered prefix extensions (Section 4). Image encoding can proceed while preserving a live 196,608-input text request. A controlled, repetitive 32K+6K workload reaches median 18.284 tokens/s at 15,626 MiB; a separate AIME 2026 evaluation scores 29/30. Coordinating execution and state lifetimes makes longer histories feasible on personal hardware.
Comments20 pages, 6 figures, 17 tables. Code, Quick Start, and reproduction materials are publicly available