arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29765cs.LGcs.NE

PruneShift:结构化剪枝中决策可靠性的评估框架

PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

Hao Ye, Gaopeng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出PruneShift框架,将结构化剪枝的预测保真度、局部保真度与决策质量分离,经多数据集实验验证,证明需分别评估剪枝的预测拟合、决策可靠性及方法质量。

中文摘要 AI 辅助

结构化剪枝采用替代目标函数,因为对所有可行掩码直接进行任务评估成本过高。大多数评估报告的是广泛采样掩码上的平均替代误差或秩相关,这些总结无法直接测试替代函数所选的掩码。我们提出PruneShift,这是一种评估框架,可将广泛的预测保真度、选择器输出附近的保真度以及所选剪枝决策的质量分离开来。我们首先证明,当归一化选择遗憾保持最大时,斯皮尔曼(Spearman)和肯德尔(Kendall)一致性可趋近于1。随后,我们基于均匀误差、选择器次优性、决策边际、密度比和比较质量推导了充分条件。该分析还产生了具有显式超额成本边界的有限池证书。四项研究对该论证的不同环节进行了测试:外部TextbookQA验证结果异质,20个同时区间中有7个支持替代函数所选掩码,6个支持其固定对比项,7个跨零;在固定的Natural Questions池中,四种设置中有一种存在严格改进;受控QQP实验在所有16个预设端点中均支持所提覆盖机制,尽管充分边界较为保守;最后,对OPT-125M进行的受限OSSCAR重构研究显示,75个主要端点中有68个的局部保真度优于广泛保真度。独立固定掩码验证在25个端点中有24个无定论,1个支持对比项。这些结果表明,预测拟合、决策可靠性和剪枝方法质量需要独立的证据支持。

英文摘要

Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average surrogate error or rank correlation on broadly sampled masks. These summaries do not directly test the mask chosen by the surrogate. We introduce PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision. We first prove that Spearman and Kendall agreement can approach one while normalized selection regret remains maximal. We then derive sufficient conditions based on uniform error, selector suboptimality, decision margin, density ratio, and comparison mass. The analysis also yields a finite pool certificate with an explicit excess cost bound. Four studies test different links in this argument. External TextbookQA confirmation is heterogeneous: 7 of 20 simultaneous intervals favor the surrogate-selected mask, 6 favor its fixed comparator, and 7 cross zero. On a fixed Natural Questions pool, strict improvement holds in one of four settings. A controlled QQP experiment supports the proposed coverage mechanism in all 16 prespecified endpoints, although the sufficient bounds are conservative. Finally, a restricted OSSCAR reconstruction study on OPT-125M shows better local than broad fidelity in 68 of 75 primary endpoints. Independent fixed-mask confirmation is inconclusive in 24 of 25 endpoints and favors the comparator in one. These results show why predictive fit, decision reliability, and pruning method quality require separate evidence.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑