剪枝在智能家居中何时失效?评估LLM在不同架构和任务复杂度下的性能退化
What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity
- Alexa Home AI, Amazon.com(亚马逊 Alexa 家庭人工智能)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究系统评估剪枝对智能家居工具调用LLM的影响,发现密集模型安全剪枝区域窄且易急剧退化,MoE模型更鲁棒,并揭示剪枝先损害上下文特异性后损害意图,强调需超越总体准确率评估剪枝效果。
AI中文摘要:
剪枝可以降低大型语言模型(LLM)的部署成本,但其对基于上下文的工具调用的影响仍知之甚少。我们系统研究了四种LLM在智能家居工具调用中的剪枝诱导退化,这些模型涵盖密集Transformer、密集混合和混合专家(MoE)架构,同时结合了深度、宽度、混合和专家剪枝方法。在剪枝后监督微调(SFT)之后,我们评估了来自三个智能家居数据集的超过19,500个实例。除了总体任务准确率外,我们沿两个维度刻画退化:动作组件(即操作、设备、参数和值)和任务复杂度。我们的结果表明,密集模型具有狭窄的安全剪枝区域,随后出现急剧退化,而MoE模型能容忍显著更多的剪枝。剪枝在损害模式级意图之前先损害基于上下文的特异性,且激进的密集剪枝可诱导系统性过度拒绝。这些发现强调了在选择剪枝后的LLM以进行可靠工具执行时,评估剪枝效果不能仅看总体准确率的重要性。
英文摘要:
Pruning can reduce the deployment cost of large language models (LLMs), but its impact on context-grounded tool calling remains poorly understood. We systematically study pruning-induced degradation in smart-home tool calling across four LLMs spanning dense Transformer, dense hybrid, and mixture-of-experts (MoE) architectures, together with depth, width, hybrid, and expert pruning methods. After post-pruning supervised fine-tuning (SFT), we evaluate more than 19,500 instances from three smart-home datasets. Beyond aggregate task accuracy, we characterize degradation along two dimensions: action components (i.e., operation, device, argument, and value) and task complexity. Our results show that dense models have narrow safe pruning regions followed by sharp degradation, while MoE models tolerate substantially more pruning. Pruning degrades grounded specificity before schema-level intent, and aggressive dense pruning can induce systematic over-refusal. These findings highlight the importance of evaluating pruning beyond aggregate accuracy when selecting pruned LLMs for reliable tool execution.