Core-KAN:基于Kolmogorov-Arnold网络的连续视觉卷积核
Core-KAN: Continuous Vision Kernels with Kolmogorov-Arnold Networks
浏览论文内容
中文总结 AI 辅助
Core-KAN是一种解耦几何尺度适配与内容依赖滤波的连续卷积算子,在少量开销下,于三类视觉任务中性能优于卷积及动态卷积核基线,为尺度自适应连续卷积提供高效通用框架。
中文摘要 AI 辅助
传统卷积核通常定义在固定离散网格上,限制了其适配异构局部结构的能力。现有自适应算子虽提升了灵活性,但常将几何尺度变化与内容依赖滤波耦合,且因逐位置生成卷积核会产生高计算成本。为解耦几何尺度适配与内容依赖滤波,同时避免昂贵的逐位置卷积核生成,本文提出连续相对尺度KAN(Core-KAN),一种相对尺度条件下的连续卷积算子。Core-KAN将输入特征映射到紧凑的潜在基空间,并用轻量尺度控制器预测相对于指数移动平均参考的局部尺度。基于KAN的生成器将深度卷积核基表示为连续坐标函数,使算子能合成任意分辨率的空间滤波器,而非局限于固定格点。它不在每个位置合成独立卷积核,而是构建一个紧凑的尺度条件卷积核响应库,并根据预测的局部尺度图对其进行插值。独立的混合控制器进一步根据局部内容组合插值后的基响应,明确将几何尺度适配与内容依赖滤波解耦。结合轻量逐点投影,该设计形成一种低秩动态卷积,其可随卷积核大小高效扩展,且易于集成到分层视觉骨干网络中。在三个代表性视觉任务上的实验表明,Core-KAN在仅产生少量参数和计算开销的情况下,始终优于强大的卷积和动态卷积核基线,为连续、尺度自适应的卷积提供了一种高效通用的框架。
英文摘要
Conventional convolutional kernels are typically defined on fixed discrete grids, limiting their ability to accommodate heterogeneous local structures. Existing adaptive operators improve flexibility but often couple geometric scale variation with content-dependent filtering, while incurring high computational cost from per-location kernel generation. To decouple geometric scale adaptation from content-dependent filtering while avoiding expensive per-location kernel generation, we propose Continuous Relative-scale KAN (Core-KAN), a relative-scale-conditioned continuous convolution operator. Core-KAN maps input features into a compact latent basis space and uses a lightweight scale controller to predict local scales relative to an exponential moving average reference. A KAN-based generator represents depth-wise kernel bases as continuous coordinate functions, allowing the operator to synthesize spatial filters at arbitrary resolutions rather than being confined to a fixed lattice. Instead of synthesizing independent kernels at every location, it constructs a compact bank of scale-conditioned kernel responses and interpolates them according to the predicted local scale map. An independent mixing controller further combines the interpolated basis responses based on local content, explicitly decoupling geometric scale adaptation from content-dependent filtering. Together with lightweight pointwise projections, this design forms a low-rank dynamic convolution that scales efficiently with kernel size and integrates readily into hierarchical vision backbones. Experiments across three representative vision tasks show Core-KAN consistently outperforms strong convolutional and dynamic-kernel baselines with only marginal parameter and computational overhead, offering an efficient, general framework for continuous, scale-adaptive convolution.