TailProp:用于视觉的内容自适应轻尾与重尾传播
TailProp: content-adaptive light- and heavy-tailed propagation for vision
- Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TailProp提出基于高斯与柯西双基传播算子的视觉主干,通过内容自适应系数在DCT域高效融合,在分类、检测、分割等任务上超越基线。
AI中文摘要:
受科学启发的视觉模型表明,显式传播动力学可以为传统令牌混合提供结构化和可解释的替代方案。然而,现有公式通常在一个特定的动力学族内构建和调整视觉传播,而视觉表示在样本、通道和网络阶段之间可能需要显著不同的空间交互。我们探索跨机制自适应传播,并引入TailProp,这是一个基于尾部传播算子(TPO)构建的分层视觉主干网络。TPO使用高斯和柯西稳定过程传播器作为互补基,分别具有快速衰减和重尾的空间影响,并预测一个内容条件化的通道级系数以自适应地组合它们。由于该系数在空间上共享,两个响应直接在DCT域中通过一对DCT/IDCT融合,对于正方形特征图(N=HW,固定通道宽度)实现O(N^1.5)的空间混合复杂度。在图像分类、目标检测、语义分割、鲁棒性和跨主干恢复任务中,TailProp始终优于匹配的传播基线;TailProp-B在ImageNet-1K上达到84.4%的Top-1准确率,在3x Mask R-CNN调度下达到50.3/44.8的框/掩码AP,在ADE20K上达到50.8%的mIoU。受控消融进一步表明,这些增益不能仅由单基传播、额外的同族分支或族内自适应阶数解释,支持互补双基传播作为视觉表示学习的有效设计原则。
英文摘要:
Science-inspired vision models show that explicit propagation dynamics can provide structured and interpretable alternatives to conventional token mixing. Existing formulations, however, typically construct and adapt visual propagation within a particular dynamical family, while visual representations can require substantially different spatial interactions across samples, channels, and network stages. We explore cross-regime adaptive propagation and introduce TailProp, a hierarchical vision backbone built upon the Tail Propagation Operator (TPO). TPO uses Gaussian and Cauchy stable-process propagators as complementary bases with rapidly decaying and heavy-tailed spatial influence, and predicts a content-conditioned channel-wise coefficient to adaptively combine them. Because this coefficient is spatially shared, the two responses are fused directly in the DCT domain with a single DCT/IDCT pair, yielding $O(N^{1.5})$ spatial mixing for square feature maps with $N=HW$ and fixed channel width. Across image classification, object detection, semantic segmentation, robustness, and cross-backbone restoration, TailProp consistently outperforms matched propagation baselines; TailProp-B reaches 84.4% Top-1 accuracy on ImageNet-1K, 50.3/44.8 box/mask AP under the 3x Mask R-CNN schedule, and 50.8% mIoU on ADE20K. Controlled ablations further show that these gains are not explained by single-basis propagation, an additional same-family branch, or within-family adaptive order alone, supporting complementary two-basis propagation as an effective design principle for visual representation learning.