发表机构
Harbin Institute of Technology, Shenzhen; Northeastern University, Shenyang; Tsinghua Shenzhen International Graduate School, Tsinghua University; Peking University Shenzhen Graduate School; Tsinghua University, Beijing(哈尔滨工业大学(深圳); 东北大学(沈阳); 清华大学深圳国际研究生院; 北京大学深圳研究生院; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出条件功能可替代性(CFS)直接刻画Transformer中的输入条件化冗余,揭示传统度量遗漏的功能关系,并证明CFS能指导动态计算,实现更优的性能-计算权衡。
AI 中文摘要
现代神经网络具有可预测的扩展规律,然而这些规律背后的机制仍不清楚。神经冗余通常通过组件重要性或表征相似性来刻画,但这两者都是间接的代理指标。我们将冗余视为一种输入条件化的动态关系:当中间计算状态引发相似的下游响应时,它们在功能上是冗余的。我们引入条件功能可替代性(CFS)来直接刻画这种功能替代关系。CFS揭示了被传统基于重要性和相似性的度量所遗漏的功能关系与缩减潜力。跨模态和Transformer家族,CFS展示了随规模扩展的系统性功能重组。受控扩展实验进一步表明,性能提升并不必然伴随可替代性的增长,而具有更独立功能结构的固定容量模型表现更好,这为收益递减提供了功能层面的解释。基于CFS的预测还能实现动态计算,其性能-计算权衡优于基于重要性的组件选择,为冗余感知计算和更高效的模型扩展指明了新方向。
英文摘要
Modern neural networks scale predictably, yet the mechanisms behind these regularities remain unclear. Neural redundancy is typically characterized by component importance or representational similarity, both indirect proxies. We view redundancy as an input-conditioned, dynamic relation: intermediate computational states are functionally redundant when they induce similar downstream responses. We introduce Conditional Functional Substitutability (CFS) to directly characterize such functional substitution. CFS exposes functional relations and reduction potential missed by conventional importance- and similarity-based measures. Across modalities and Transformer families, CFS reveals systematic functional reorganization with scale. Controlled scaling further shows that performance gains need not track growth in substitutability, while fixed-capacity models with more independent functional structure perform better, providing a functional account of diminishing returns. Predicted CFS further enables dynamic computation with a better performance--computation trade-off than importance-based component selection, suggesting new directions for redundancy-aware computation and more efficient model scaling.
Comments21 pages, 4 figures, 7 tables