发表机构
Montclair State University; New Jersey Institute of Technology(蒙特克莱尔州立大学; 新泽西理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出压缩损伤决定校准数据需求,量化影响小,剪枝损伤大时用金融任务示例校准可恢复性能,建议先测损伤再定数据。
AI 中文摘要
后训练量化和剪枝依赖于一个小型校准语料库。金融等专业领域是否需要领域匹配的校准数据仍无定论。我们认为,答案取决于压缩引起的任务级损伤,而非领域不匹配。如果压缩保留了目标能力,改变校准语料库影响甚微。如果压缩导致较大损失,任务格式化的校准可以恢复部分损失。我们在两个模型家族、六种压缩配置、三个令牌匹配的校准语料库以及十个金融分类和数值问答任务上检验了这一假设。结果支持该假设。量化在很大程度上保留了任务性能,此时校准选择影响甚微。剪枝将数值问答准确率降低了超过40个百分点。在这些受损设置中,另一个通用语料库无济于事,而FinMix(金融任务示例的混合)恢复了大部分损失。损伤与恢复之间的联系在模型家族和规模上均成立。这些发现支持一条实用规则:首先衡量任务特定的压缩损伤,仅在损伤较大时构建专门的校准数据。
英文摘要
Post-training quantization and pruning rely on a small calibration corpus. Whether specialized domains such as finance require domain-matched calibration data remains unsettled. We argue that the answer depends on the task-level damage caused by compression rather than on domain mismatch. If compression preserves the target capability, changing the calibration corpus has little effect. If compression causes large losses, task-formatted calibration can recover part of the loss. We test this hypothesis across two model families, six compression configurations, three token-matched calibration corpora, and ten financial classification and numerical question-answering tasks. The results support this hypothesis. Quantization largely preserves task performance, and calibration choice has little effect in this case. Pruning reduces numerical QA accuracy by over 40 points. In these damaged settings, another generic corpus does not help, while FinMix, a mixture of financial task examples, recovers a large part of the loss. The link between damage and recovery holds across model families and scales. These findings support a practical rule. Measure task-specific compression damage first, and construct specialized calibration data only when the damage is large.