arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29592cs.CVcs.AI

QINA:用于预训练视觉模型的量子启发非线性适配器

QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models

Mostafa Mehdipour Ghazi

AI总结:

针对预训练视觉模型在冻结骨干下的适配难题,提出量子启发非线性适配器(QINA),通过可学习三角特征提升与有界聚合实现谱重塑,实验证明其优于多种基线,无需量子硬件。

AI中文摘要:

在有限数据和冻结骨干网络约束下,适配大型预训练视觉模型仍然是迁移学习中的核心挑战。尽管轻量级适配器和参数高效微调方法被广泛采用,但大多数依赖通用多层感知机或低秩线性更新,对特征变换的谱和几何结构的控制有限。我们研究结构化非线性特征提升是否能在冻结机制中改善表示对齐。我们引入量子启发非线性适配器(QINA),这是一种紧凑模块,执行可学习的三角特征提升,随后进行有界非线性聚合。该设计引入具有显式范数相关Lipschitz界的有结构振荡基函数,能够在不增加感受野或显著扩大参数数量的情况下,对预训练表示进行谱重塑。重要的是,该方法完全在标准深度学习框架内运行,不需要量子硬件。通过在自然和医学图像数据集、分类和分割任务、多种适配器和放置位置以及不同训练预算上的系统实验,我们表明冻结机制中的性能主要受表示限制。非线性提升改善了适配,所提出的结构化三角公式始终优于恒等基线、固定傅里叶特征映射和参数匹配的基线适配器。在评估的冻结骨干设置中,结构化谱参数化提供了比通用非线性适配器更有效的归纳偏置。这项工作强调了针对大型预训练视觉模型的几何和谱感知适配机制的重要性。

英文摘要:

Adapting large pretrained vision models under limited data and frozen-backbone constraints remains a central challenge in transfer learning. While lightweight adapters and parameter-efficient fine-tuning methods are widely adopted, most rely on generic multilayer perceptrons or low-rank linear updates, offering limited control over the spectral and geometric structure of feature transformations. We investigate whether structured nonlinear feature lifting can improve representational alignment in frozen regimes. We introduce Quantum-Inspired Nonlinear Adapters (QINA), compact modules that perform learnable trigonometric feature lifting followed by bounded nonlinear aggregation. The design induces structured oscillatory basis functions with an explicit norm-dependent Lipschitz bound, enabling spectral reshaping of pretrained representations without increasing the receptive field or significantly expanding parameter count. Importantly, the method operates entirely within standard deep learning frameworks and does not require quantum hardware. Through systematic experiments across natural and medical imaging datasets, classification and segmentation tasks, multiple adapters and placements, and varying training budgets, we show that performance in frozen regimes is primarily representation-limited. Nonlinear lifting improves adaptation, and the proposed structured trigonometric formulation consistently outperforms identity baselines, fixed Fourier feature mappings, and parameter-matched baseline adapters. Within the evaluated frozen-backbone settings, structured spectral parameterization provides a more effective inductive bias than generic nonlinear adapters. This work highlights the importance of geometry- and spectrum-aware adaptation mechanisms for large pretrained vision models.

补充信息

↑