发表机构
University of Arizona; University of Pittsburgh; University of Central Florida(亚利桑那大学; 匹兹堡大学; 中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出轨迹锚定剪枝(TAP),首个针对RL训练智能体的结构化剪枝框架,通过锚定教师轨迹的在线策略恢复和迭代通道选择,在移除60%FFN通道时保持ALFWorld和WebShop上99.2%和88.0%的成功率,并显著降低GPU时间。
AI 中文摘要
新兴的长时程智能体任务需要重复调用模型,这加剧了本已昂贵的语言模型的推理成本。尽管狭窄的智能体任务表明在不损失性能的情况下进行激进模型剪枝的潜力,但实证结果显示,为问答任务提出的现有方法在应用于智能体模型时会严重降低任务性能。我们将这一失败追溯到两个决策:剪枝什么以及如何恢复。对于剪枝,一次性重要性估计无法跟踪剪枝模型的适应过程。对于恢复,离线蒸馏仅覆盖教师前缀,而全轨迹在线策略蒸馏会导致学生错误在回合间累积。在这项工作中,我们提出了轨迹锚定剪枝(TAP),这是首个针对强化学习(RL)训练的智能体的结构化剪枝框架。TAP将结构化剪枝与高效的在线策略恢复相结合,将交互锚定到教师轨迹,同时允许学生生成每个推理-行动响应。冻结的稠密教师监督学生的响应前缀,解决了响应内训练-推理不匹配问题,同时防止学生引发的偏差在训练回合间传播。TAP不是一次性剪枝,而是使用恢复目标在恢复的学生上的梯度重新对通道进行评分,将迭代通道选择与不断演化的策略联系起来。在移除60%的前馈网络通道后,TAP在ALFWorld和WebShop上分别保留了稠密7B智能体任务成功率的99.2%和88.0%,同时每个成功任务的GPU时间分别减少了约22%和17%。这些结果证明了在有限恢复预算下对长时程智能体进行有效结构化压缩的可行性。
英文摘要
Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, empirical results show existing methods proposed for question answering tasks severely degrade task performance when applied to agentic models. We trace this failure to two decisions: what to prune and how to recover. For pruning, one-shot importance estimates fail to track how the pruned model adapts. For recovery, offline distillation covers only teacher prefixes, while full-trajectory on-policy distillation causes student errors to compound across turns. In this work, we propose Trajectory-Anchored Pruning (TAP), the first structural pruning framework for reinforcement learning (RL)-trained agents. TAP couples structural pruning with efficient on-policy recovery, anchoring interactions to teacher trajectories while allowing the student to generate each reasoning-action response. A frozen dense teacher supervises the student's response prefixes, addressing within-response training-inference mismatch while preventing student-induced deviations from propagating across training turns. Instead of one-shot pruning, TAP re-scores channels using gradients of the recovery objective on the recovered student, connecting iterative channel selection to the evolving policy. With 60% of FFN channels removed, TAP retains 99.2% and 88.0% of the dense 7B agents' task success rates on ALFWorld and WebShop, respectively, while reducing GPU time per successful task by approximately 22% and 17%. These results demonstrate effective structural compression of long-horizon agents under a limited recovery budget.