arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

C$^2$Path:用于视觉语言增量目标检测的类条件路径解耦

C$^2$Path: Class-Conditional Pathway Decoupling for Vision-Language Incremental Object Detection

Lecheng Xu, Feifei Shao, Ouyangzi Ye, Zhen Wang, Lin Li, Kexin Li, Zhao Wang, Changqin Huang

arXiv 2608.21937首次发表:更新:

发表机构

Zhejiang University; The Hong Kong University of Science and Technology; Zhejiang Tobacco Monopoly Administration(浙江大学; 香港科技大学; 浙江省烟草专卖局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言增量目标检测的类知识耦合问题,提出类条件路径解耦框架C$^2$Path,通过类别专家库与类条件解耦模块实现隔离更新,在COCO 2017上性能优于现有方法

AI 中文摘要

增量目标检测(IOD)旨在让检测器在持续学习新类别的同时保留已获取的知识,但现有方法存在两类类知识耦合问题:共享参数更新导致的类边界侵蚀,以及混合特征编码引发的类表示纠缠。我们认为,有效的增量学习需要类特定计算路径,以实现隔离的参数更新和分离的类注入。为此,我们提出C$^2$Path,这是一种用于视觉语言增量目标检测的类条件路径解耦框架,利用令牌级类线索为不同类别建立专用且可更新的计算路径。具体而言,C$^2$Path引入类别专家库和类条件解耦模块:专家库由可学习的低秩计算节点构成,用于捕获类别特定知识;解耦模块生成类感知路由信号,动态从这些专家中组合ClassLoRA适配器,从而形成类特定计算路径,实现跨类别的隔离更新与分离注入。在COCO 2017数据集上的大量实验(涵盖多种增量学习设置)表明,C$^2$Path始终优于现有最先进方法,为视觉语言检测器的持续类别扩展提供了有效且可扩展的解决方案。

英文摘要

Incremental Object Detection (IOD) aims to enable detectors to continuously learn novel categories while preserving previously acquired knowledge. However, existing methods suffer from two forms of \textbf{class knowledge coupling}: class boundary erosion induced by shared parameter updates and class representation entanglement arising from mixed feature encoding. We argue that effective incremental learning requires class-specific computational pathways that enable isolated parameter updates and separated class-wise injection. To this end, we propose \textbf{C$^2$Path}, a class-conditional pathway decoupling framework for vision-language incremental object detection that leverages token-level class cues to establish dedicated and updatable computational pathways for different categories. Specifically, C$^2$Path introduces a category expert library and a class-conditional decoupling module. The expert library consists of learnable low-rank computational nodes that capture category-specific knowledge, while the decoupling module generates class-aware routing signals to dynamically compose \textit{ClassLoRA} adapters from these experts, thereby forming class-specific computational pathways for isolated updates and separated injection across categories. Extensive experiments on COCO 2017 under multiple incremental learning settings demonstrate that C$^2$Path consistently outperforms state-of-the-art methods, providing an effective and scalable solution for continual category expansion in vision-language detectors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑