彩票票并非部署票
Lottery Tickets Are Not Deployment Tickets
浏览论文内容
中文总结 AI 辅助
该研究在部署层面发现,虽彩票票等稀疏模型可匹配密集模型准确率,但行为仍有差异,替换会改变7%-10%的决策,凸显干净准确率认证的局限性,明确稀疏模型替换现有模型的实际挑战。
中文摘要 AI 辅助
现有文献中关于稀疏化、压缩和彩票票(Lottery Tickets,LTs)如何改变模型行为的研究结果不一,部分研究观察到有益效果,部分则发现不利影响。此外,过往研究未考虑实际部署场景,在这些场景中,现有模型的决策逻辑已固定。为从实际角度评估这些混合发现,我们在部署层面研究生产替换问题,即准确率匹配的彩票票或其他稀疏候选模型能否在不重新配置下游决策逻辑的情况下替换现有密集模型。因此,我们审计了广泛的、与协议相关的部署相关行为,包括校准、分布外(OOD)响应、类级可靠性、表征和下游策略决策,并总结了排除干净准确率后的偏差,即行为兼容性距离。在大量实验中,稀疏候选模型反复恢复密集参考模型的准确率,但行为仍存在差异;在多个研究带匹配设置中,彩票票还表现出更低的损坏准确率。在固定阈值策略诊断的小差距设置中,彩票票替换会改变7%至10%的接受-审核决策。这种变动恰好造成了直接替换本应避免的负担:重新配置和验证下游决策逻辑。这些发现确立了干净准确率认证的局限性:与固定现有模型的兼容性建立,不同于将变动唯一归因于稀疏性,或把每个测量偏差都视为有害。我们的理论解释了路由结果:即使逐点精确的Top-1一致性也无法约束固定阈值决策变动,且操作边界附近的小置信度偏移会产生一阶路由变动。
英文摘要
Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others. Moreover, prior work has not considered actual deployment conditions, where decision logic is already fixed for the incumbent. To assess these mixed findings from a practical standpoint, we study the production-replacement question at the deployment level, namely whether an accuracy-matched lottery ticket or another sparse challenger can replace an incumbent dense model without reconfiguring downstream decision logic. We therefore audit a broad, protocol-specific panel of deployment-relevant behaviors spanning calibration, OOD response, class-level reliability, representations, and downstream policy decisions, and summarize clean-accuracy-excluded deviations with a behavioral-compatibility distance. Across extensive experiments, sparse candidates repeatedly recover dense-reference accuracy yet remain behaviorally different; in several study-band-matched settings, LTs also show lower corruption accuracy. In small-gap settings with fixed-threshold policy diagnostics, lottery-ticket replacement changes 7% to 10% of accept--review decisions. This churn creates precisely the burden that drop-in replacement is meant to avoid: reconfiguring and revalidating downstream decision logic. These findings establish the limits of clean-accuracy certification: Establishing compatibility with a fixed incumbent is distinct from attributing churn uniquely to sparsity or treating every measured deviation as harmful. Our theory explains the routing result: Even exact pointwise top-1 agreement cannot bound fixed-threshold decision changes, and small confidence shifts near the operating boundary can generate first-order routing churn.