发表机构
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CrossMambaTuning将状态空间模型与跨层交互机制结合,通过Mamba适配器和SICA实现参数高效微调,在多项机器视觉任务上达SOTA性能,参数开销降低72%。
AI 中文摘要
为降低部署成本与重训练开销,将预训练的学习型图像压缩(LIC)模型适配至下游机器视觉任务已受到越来越多关注。然而,现有方法通常将微调模块独立插入冻结的骨干网络,缺乏显式的跨层协调机制。为解决这一局限,我们提出名为CrossMambaTuning的新型框架,该框架将状态空间模型与跨层交互机制相结合,用于参数高效微调。具体而言,我们设计了一种高效的Mamba适配器,配备任务特定提示与多尺度分支,以精准捕获局部特征与全局依赖。此外,我们引入了利用参数共享策略的尺度不变跨层适配器(SICA),用于融合不同尺度的任务信息并减少冗余。大量实验表明,CrossMambaTuning在多项机器视觉任务上实现了最优(SOTA)性能,与SOTA方法相比参数开销降低了72%。代码可在指定URL获取。
英文摘要
To reduce deployment cost and retraining overhead, adapting pretrained learned image compression (LIC) models to downstream machine vision tasks has attracted growing attention. However, existing methods typically insert fine-tuning modules independently into frozen backbones, lacking explicit mechanisms for cross-layer coordination. To address this limitation, we propose a novel framework named CrossMambaTuning, which integrates State Space Models with cross-layer interaction mechanisms for parameter-efficient fine-tuning. Specifically, we design an efficient Mamba adapter equipped with task-specific prompts and multi-scale branching to precisely capture both local features and global dependencies. Furthermore, we introduce a Scale-Invariant Cross-Layer Adapter (SICA) utilizing a parameter-sharing strategy to fuse task information across different scales and reduce redundancy. Extensive experiments demonstrate that CrossMambaTuning achieves state-of-the-art (SOTA) performance on multiple machine vision tasks, reducing parameter overhead by 72\% compared to SOTA methods. Code is available at https://github.com/rsr1123/CrossMambaTuning.
CommentsAccepted by ACM MM26