arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15498cs.LGmath.DS

相同路径,不同机制:循环与堆叠Transformer编码器在12导联心电图上的机理比较

Same path, different: a mechanistic comparison of looped and stacked transformer encoders on 12-lead ECG

  • Samsung AI Center(三星人工智能中心)
  • IFTR, Polish Academy of Sciences(波兰科学院基础技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Pawel Olszowiec, Michal Byra, Grzegorz Gruszczynski, Grzegorz Stefanski, Alberto Presta

AI总结:

本研究以bViT为例,对比循环与堆叠Transformer在12导联ECG分类中的表征与动态差异,发现两者表征相似但动态不同,循环架构参数效率高且行为更稳健。

AI中文摘要:

循环Transformer通过复用权重而非堆叠$L$个不同层,因其参数效率而正被广泛采用[1,2,3]。然而,循环架构与堆叠架构之间精确的表征和动态差异仍未得到刻画。本文以bViT模型[1]为例进行了一项受控研究,该模型将一个权重共享的块应用$L$次。我们在相同的训练协议下,针对PTB-XL数据集上的12导联心电图(ECG)分类任务训练了两个模型:bViT和标准ViT[4]。尽管参数减少了$8.9\ imes$,bViT仍达到了与ViT相当的准确率。几何相似性度量表明,两种架构以等效的规范顺序构建了可比较的潜在表征。关键在于,它们的动态行为不同:bViT表现出更小的步长和患者间敏感性,以及在数据流形之外的近中性行为,而ViT则表现出表征维度坍缩和分布外特征扩张。

英文摘要:

Recurrent Transformers reusing their weights rather than stacking $L$ distinct layers are becoming widely adopted due to their parameter efficiency [1,2,3]. However, the exact representational and dynamical differences between looped and stacked architectures remain uncharacterized. This paper presents a controlled study on the example of bViT model [1] applying one weight-tied block $L$ times. We train two models: bViT and standard ViT [4] on 12-lead electrocardiogram (ECG) classification tasks from the PTB-XL dataset under identical training protocols. Despite an $8.9\times$ parameter reduction, bViT achieves accuracy parity with ViT. Geometric similarity metrics demonstrate that both architectures construct comparable latent representations in an equivalent canonical order. Crucially, their dynamics differ: bViT exhibits smaller step sizes and inter-patient sensitivity, as well as near-neutral behavior away from the data manifold, whereas ViT exhibits collapsing dimensionality of representations and out-of-distribution feature expansion.

补充信息

↑