arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrossMambaTuning:面向机器视觉压缩的协同空间与跨层适配

CrossMambaTuning: Synergistic Spatial and Cross-Layer Adaptation for Machine Vision Compression

Haobo Xiong, Shaobo Liu, Kai Liu, Chongyang Ding

arXiv 2608.25568首次发表:更新:

发表机构

School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CrossMambaTuning将状态空间模型与跨层交互机制结合,通过Mamba适配器和SICA实现参数高效微调,在多项机器视觉任务上达SOTA性能,参数开销降低72%。

AI 中文摘要

为降低部署成本与重训练开销,将预训练的学习型图像压缩(LIC)模型适配至下游机器视觉任务已受到越来越多关注。然而,现有方法通常将微调模块独立插入冻结的骨干网络,缺乏显式的跨层协调机制。为解决这一局限,我们提出名为CrossMambaTuning的新型框架,该框架将状态空间模型与跨层交互机制相结合,用于参数高效微调。具体而言,我们设计了一种高效的Mamba适配器,配备任务特定提示与多尺度分支,以精准捕获局部特征与全局依赖。此外,我们引入了利用参数共享策略的尺度不变跨层适配器(SICA),用于融合不同尺度的任务信息并减少冗余。大量实验表明,CrossMambaTuning在多项机器视觉任务上实现了最优(SOTA)性能,与SOTA方法相比参数开销降低了72%。代码可在指定URL获取。

英文摘要

To reduce deployment cost and retraining overhead, adapting pretrained learned image compression (LIC) models to downstream machine vision tasks has attracted growing attention. However, existing methods typically insert fine-tuning modules independently into frozen backbones, lacking explicit mechanisms for cross-layer coordination. To address this limitation, we propose a novel framework named CrossMambaTuning, which integrates State Space Models with cross-layer interaction mechanisms for parameter-efficient fine-tuning. Specifically, we design an efficient Mamba adapter equipped with task-specific prompts and multi-scale branching to precisely capture both local features and global dependencies. Furthermore, we introduce a Scale-Invariant Cross-Layer Adapter (SICA) utilizing a parameter-sharing strategy to fuse task information across different scales and reduce redundancy. Extensive experiments demonstrate that CrossMambaTuning achieves state-of-the-art (SOTA) performance on multiple machine vision tasks, reducing parameter overhead by 72\% compared to SOTA methods. Code is available at https://github.com/rsr1123/CrossMambaTuning.

CommentsAccepted by ACM MM26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑