量化智能体剖析:代码合成智能体工作负载中的VRAM稳定性与预测
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
浏览论文内容
中文总结 AI 辅助
该研究针对基于LangGraph的AgentK智能体,在1920条轨迹上评估量化LLM的VRAM消耗,提出闭式峰值内存预测模型,发现编译成功率受LLM容量限制,且无需复杂VRAM预测模型。
中文摘要 AI 辅助
针对大语言模型(LLM)推理的峰值VRAM消耗分析模型,将内存分解为权重存储、KV缓存和激活项,这些项由步数、工具调用和上下文扩展参数化。我们在严格限定范围的测量研究中对该分解进行了实证评估:基于LangGraph的CUDA内核合成智能体(AgentK)、4位量化系列(Q4 K M)、单张NVIDIA H100 GPU,以及四个LLM骨干模型,共1920条轨迹。聚焦峰值内存预测行为,我们报告两个主要观察结果。第一,当提供两个经验常数(加载权重VRAM和固定激活内存开销)时,闭式分析模型可达到有竞争力的准确率。在提供实时GPU读数和真实轨迹参数的情况下,该闭式模型在四个骨干模型中的三个上表现与最佳学习基线相当或更优(测试平均绝对百分比误差MAPE为2.2%-4.4%,而学习基线为3.4%-6.5%,p值为0.76)。例外情况是最小的骨干模型Phi-4-mini,其VRAM方差极小(变异系数CV为0.3%),导致动态建模的表现不如简单回归。第二,编译成功率按骨干模型容量严格分化(从Phi-4-mini的5.7%到Qwen2.5-Coder-14B的62.0%),表明功能性代码合成仍受LLM内在能力而非可用内存的限制。此外,由于所有骨干模型的总体峰值内存方差极低(CV为0.3%-9.4%),学习提示特征回归相比常数均值基线的提升在统计上不显著。因此,我们发现没有理由在高度量化、权重主导的场景中部署复杂的预测VRAM模型。我们发布了评估语料库和匿名框架以支持复现。
英文摘要
Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and context expansion. We evaluate this decomposition empirically within a strictly scoped measurement study: a LangGraph-based CUDA-kernel-synthesis agent (AgentK), a 4-bit quantization family (Q4 K M), a single NVIDIA H100 GPU, and four LLM backbones across 1,920 trajectories. Focusing on peak-memory forecasting behavior, we report two primary observations. First, closed-form analytical models achieve competitive accuracy when provided with two empirical constants: loaded-weight VRAM and a fixed activation-memory overhead. Supplied with live GPU readings and ground-truth trajectory parameters, the closed-form model matches or outperforms the best learned baseline on three of the four backbones (test MAPE 2.2-4.4% vs. 3.4-6.5%, p = 0.76). The exception is the smallest backbone (Phi-4-mini), where minimal VRAM variance (CV 0.3%) causes dynamic modeling to underperform simple regression. Second, compile success strictly bifurcates by backbone capacity (from 5.7% for Phi-4-mini to 62.0% for Qwen2.5-Coder-14B), demonstrating that functional code synthesis remains constrained by intrinsic LLM capabilities rather than available memory. Furthermore, because overall peak-memory variance is remarkably low across all backbones (CV 0.3-9.4%), learned prompt-feature regression offers statistically insignificant improvements over a constant-mean baseline. Consequently, we find no justification for deploying complex predictive VRAM models in highly quantized, weight-dominated regimes. We release the evaluated corpus and anonymized framework to support replication.
发表机构
- Nokia Germany(诺基亚德国)
机构由 AI 辅助整理,请以论文原文为准。