Lean-SAM2:用于SAM2的目标锚定内存与编码器加速
Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2
浏览论文内容
中文总结 AI 辅助
针对SAM2内存开销大、特征提取冗余及复杂场景性能差的问题,提出Lean-SAM2,集成目标锚定内存修剪、时间压缩、风险感知路由三种机制,在多基准测试中平衡了准确性和效率,实现推理加速并提升分数。
中文摘要 AI 辅助
段分割模型2(SAM2)推进了时间可提示分割,但由于大量内存交叉注意力开销和冗余全帧视觉特征提取,其部署仍受阻碍。近期方法通过启发式内存修剪和基于窗口的稀疏路由探索效率,但在复杂分割场景中性能会严重下降。为此提出Lean-SAM2,它集成了三种协作机制:目标锚定内存修剪、带保险内存的时间压缩、目标锚定风险感知路由。大量评估表明Lean-SAM2在准确性和效率之间实现了更好的平衡。例如,在LVOSv2验证数据集上,Lean-SAM2在SAM2.1-Large和SAM2.1-Base+上分别实现了1.412倍和1.417倍的整体推理加速,显著优于Efficient-SAM2,同时将相应的J&F分数提高了5.0%和3.6%。
英文摘要
The Segment Anything Model 2 (SAM2) has advanced temporal promptable segmentation, yet its deployment remains hindered by heavy memory cross-attention overhead and redundant full-frame visual feature extraction. While recent methods explore efficiency via heuristic memory pruning and window-based sparse routing, they typically suffer from catastrophic performance degradation in complex segmentation scenarios replete with occlusions and distractors. To resolve these limitations, we propose \textbf{Lean-SAM2}, a holistic lightweight framework designed to address the above vulnerabilities while systematically eliminating computational redundancies. Specifically, Lean-SAM2 integrates three collaborative mechanisms: (1) Target-Anchored Memory Pruning (TAMP) safeguards target tokens against deceptive attention by modulating raw attention significance with semantic consistency against prompt-derived foreground anchors; (2) Temporal Condensation with Insurance Memory (TCIM) condenses historical context via a visibility-gated fusion while conditionally archiving high-confidence entries in a parallel insurance bank; and (3) Target-Anchored Risk-Aware Routing (TARR) selectively activates the heavy image encoder for target-related windows based on anchor similarity, utilizing a risk-aware fallback policy to trigger full-frame refreshes during volatile transitions. Extensive evaluations across multiple challenging benchmarks demonstrate that Lean-SAM2 establishes a superior balance between accuracy and efficiency. For example, on the LVOSv2 validation dataset, Lean-SAM2 achieves overall inference speedups of $1.412\times$ and $1.417\times$ on the SAM2.1-Large and SAM2.1-Base+, respectively, significantly outperforming Efficient-SAM2 while boosting the corresponding $\mathcal{J}\&\mathcal{F}$ scores by $5.0\%$ and $3.6\%$. Code is available at https://github.com/DeawhaleQwQ/Lean-SAM2.
发表机构
- School of Biomedical Engineering, Hainan University(海南大学生物医学工程学院)
- Department of Electronics and Electrical Engineering, Keio University(庆应义塾大学电子电气工程系)
- College of Computer Science, Shenyang Aerospace University(沈阳航空航天大学计算机科学学院)
- School of Computer Science and Technology, Hainan University(海南大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。