APQF:基于智能体分析引导的结构化剪枝与混合精度量化及自适应微调
APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
浏览论文内容
中文总结 AI 辅助
APQF是结合LLM引导与分析支撑的自动化剪枝量化框架,在ImageNet等数据集上实现高压缩比且精度接近基线,优于现有联合剪枝量化方法。
中文摘要 AI 辅助
现代深度神经网络性能强劲,但规模庞大导致成本高昂、运行缓慢,尤其在资源受限的边缘设备上表现突出。剪枝与量化可解决该问题,但依赖人工专家选择,且算法难以跨架构应用;统一设置未考虑各层对压缩的不同响应,会造成精度损失。本文提出APQF,一种智能体分析引导的框架,将结构化剪枝、混合精度量化感知训练与精度恢复整合为一条自动化流水线。分析智能体测量模型的成本分布及各部分对剪枝的敏感度,该证据驱动LLM规划器提出各层剪枝率、各层层位宽及恢复策略,且所有方案在执行前均经过验证。据我们所知,APQF是首个将LLM引导、分析支撑的决策与全训练感知剪枝量化流水线结合的框架,适用于卷积神经网络(CNNs)和视觉Transformer(ViTs)。我们在ResNet、VGG7、ViT、DeiT和Swin模型上,使用ImageNet-1k和CIFAR-10数据集评估APQF。在ImageNet上,它将计算量降至原始位操作的5.6%-7.7%,即13-18倍压缩,同时精度接近基线;在20万张图像的预算下,其Top-1精度比现有联合剪枝量化方法高出约17个百分点。在CIFAR-10上,它在5种架构中的4种上实现比该方法更优的压缩效果。在VGG7上,它仅使用基线位操作的0.41%就达到93.15%的精度,是该压缩水平下唯一优于全精度基线的方法。消融实验显示,在相同计算量下,统一压缩损失的精度最高;若规划器缺少分析数据,所有模型的性能均会受损。6种LLM规划器(包括免费开放权重的)在Swin-Tiny上均达到97.4%-97.9%的精度。
英文摘要
Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning and quantization address this, but rely on manual, expert choices and on algorithms that are hard to apply across architectures. Uniform settings also ignore how differently individual layers respond to compression, which costs accuracy. We introduce APQF, an agentic profiling-guided framework that combines structured pruning, mixed-precision quantization-aware training, and accuracy recovery in one automated pipeline. A profiling agent measures how cost is distributed across the model and how sensitive each part is to pruning, and this evidence drives per-layer pruning ratios, per-layer bit-widths, and the recovery strategy, all proposed by LLM planners and validated before execution. To our knowledge, APQF is the first framework to combine LLM-guided, profiling-grounded decisions with a fully training-aware pruning and quantization pipeline for both CNNs and vision transformers. We evaluate APQF on ResNet, VGG7, ViT, DeiT, and Swin using ImageNet-1k and CIFAR-10. On ImageNet it cuts compute to 5.6-7.7 percent of the original bit-operations, a 13-18x reduction, while keeping accuracy close to the baseline, and under a 200K-image budget it stays roughly 17 points higher in Top-1 than existing joint pruning and quantization methods. On CIFAR-10 it compresses further than that method on four of five architectures. On VGG7 it reaches 93.15 percent using only 0.41 percent of baseline bit-operations, the only method at that compression level to improve on its full-precision baseline. Ablations show that uniform compression loses the most accuracy at matched compute, and that withholding profiling data from the planner hurts every model. Six LLM planners, including free open-weight ones, all reach 97.4-97.9 percent on Swin-Tiny.
发表机构
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。