arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

没有免费的效率:重新审视训练效率与模型脆弱性之间的权衡

No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability

Yiyong Liu, Jun Sakuma, Michael Backes, Rui Wen

arXiv 2609.33898首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security; Institute of Science Tokyo(CISPA亥姆霍兹信息安全中心; 东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统调查训练效率与模型脆弱性之间的权衡,发现高效训练策略增加对抗与隐私攻击敏感性,并呼吁多目标训练以兼顾性能、成本与安全。

AI 中文摘要

训练效率已成为近期基础模型进展的核心驱动力。为了克服大规模训练中巨大的计算和数据需求,研究人员越来越多地采用选择性数据采样、高效预训练和简化强化学习流程等策略。虽然这些策略大幅降低了开销,但它们引发了一个关键但被忽视的问题:效率是否以模型鲁棒性和安全性为代价?据我们所知,我们首次对效率-脆弱性权衡进行了系统的跨领域调查。在视觉和语言模型中,我们表明面向效率的训练增加了对对抗性攻击和隐私攻击的敏感性。我们通过分析模型的内部几何结构和功能表征来刻画这种脆弱性,证明所评估的高效变体一致地表现出更尖锐的损失几何形状以及表征结构的系统性变化。我们进一步将分析扩展到“零强化学习训练”,发现使用简化强化学习配方训练的模型比通过传统对齐流程训练的模型表现出显著更大的灾难性遗忘倾向和更明显的过度自信。我们的发现表明,训练效率很少是“免费的午餐”;相反,最小化计算的机制可能会无意中损害安全性。我们最后呼吁向多目标训练范式转变,该范式同时优化性能、成本和安全性。

英文摘要

Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient pre-training, and simplified reinforcement learning pipelines. While these strategies drastically reduce overhead, they prompt a critical, yet neglected question: Is efficiency achieved at the expense of model robustness and security? To our knowledge, we present the first systematic cross-domain investigation of the efficiency-vulnerability trade-off. Across vision and language models, we show that efficiency-oriented training increases susceptibility to adversarial and privacy attacks. We characterize this vulnerability by analyzing the models' internal geometry and functional representations, demonstrating that the evaluated efficient variants consistently exhibit sharper loss geometry together with systematic changes in representational structure. We further extend our analysis to "zero RL training", finding that models trained using simplified RL recipes exhibit substantially greater susceptibility to catastrophic forgetting and more pronounced overconfidence than those trained through conventional alignment pipelines. Our findings suggest that training efficiency is rarely a "free lunch"; rather, the mechanisms that minimize computation can inadvertently compromise safety. We conclude by calling for a paradigm shift toward multi-objective training that jointly optimizes for performance, cost, and security.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑