arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态参数化并非动态推理

Dynamic Parameterization Is Not Dynamic Inference

Zongfei Li, Yuan-yih Shang, Guozhong Luo

arXiv 2607.26192首次发表:更新:

AI 中文总结

该研究提出冻结控制器审计(FCA)方法,发现动态参数化不代表动态推理,功能动态不代表计算节省,动态模型相关主张需单独报告系数变化等关键指标。

AI 中文摘要

依赖输入的控制器系数常被视为动态推理或计算节省的证据。这种解释混淆了三个属性:系数变化、冻结模型对系数分配给输入方式的依赖性,以及条件执行。我们聚焦于第二个属性,提出了冻结控制器审计的一般原则。我们提供了一个具体实现:冻结控制器审计(FCA),该方法在未受扰动的轨迹上缓存完整的系数张量,禁用控制器,并通过跨输入重新分配、令牌洗牌以及从独立校准集估计的静态配置文件来重放冻结模型。由于系数在任何干预前已被缓存,重放下的性能变化可衡量分配依赖性,而无需通过在受扰动的隐藏状态上重新计算控制器来反馈。在七个独立训练的76M参数的FeatureGate Transformer模型和三个504M参数的模型上,静态层配置文件分别保留了98.70%和99.43%的“正确到全局均值”性能差距。层身份解释了87%至96%的系数方差。FeatureGate仍会执行每个Transformer块,其测得的推理速度比Dense慢30.8%。在公开的MUDDPythia-1.4B检查点上,跨输入重新分配和令牌洗牌分别使负对数似然(NLL)增加1.9067和2.9637。这些惩罚表明,模型强烈依赖于内容条件下的跨层分配,MUDDPythia同样会执行每个Transformer块。结果表明,仅动态参数化无法确立动态推理,功能动态也无法确立计算节省。关于动态模型的主张应单独报告系数变化、冻结模型的功能依赖性以及实际执行情况。

英文摘要

Input-dependent controller coefficients are often treated as evidence of dynamic inference or computational savings. This interpretation conflates three properties: coefficient variation, dependence of a frozen model on how coefficients are assigned to inputs, and conditional execution. We focus on the second property and formulate a general principle of frozen-controller auditing. We provide one concrete implementation, Frozen-Controller Auditing (FCA), which caches the complete coefficient tensor along an unperturbed trajectory, disables the controller, and replays the frozen model with cross-input reassignment, token shuffling, and static profiles estimated from an independent calibration set. Because the coefficients are cached before any intervention, performance changes under replay measure assignment dependence without feedback from recomputing the controller on perturbed hidden states. Across seven independently trained 76M FeatureGate Transformers and three 504M models, static layerwise profiles retain 98.70% and 99.43% of the Correct-to-GlobalMean performance gap, respectively. Layer identity explains 87% to 96% of the coefficient variance. FeatureGate nevertheless executes every Transformer block, and its measured inference is 30.8% slower than Dense. On the public MUDDPythia-1.4B checkpoint, cross-input reassignment and token shuffling increase NLL by 1.9067 and 2.9637, respectively. These penalties show that the model depends strongly on content-conditioned cross-layer assignment. MUDDPythia also executes every Transformer block. The results show that dynamic parameterization alone does not establish dynamic inference and that functional dynamics do not establish computational savings. Claims about dynamic models should separately report coefficient variation, functional dependence of the frozen model, and actual execution.

Comments10 pages, 5 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑