A2DINOv3:通过社会化协作重新思考多模态目标检测
A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
浏览论文内容
中文总结 AI 辅助
本文提出A2DINOv3框架,通过社会化协作协议让RGB与红外分支以异构专家模式选择性交互,结合零初始化策略,在四个多模态基准上实现了多模态目标检测的SOTA性能。
中文摘要 AI 辅助
多模态目标检测对于在低光照、恶劣环境等挑战性条件下实现鲁棒的场景理解至关重要。近期的视觉基础模型(如DINOv3)展现出强大的表征能力,但将其适配到多模态场景仍具挑战性。现有密集跨模态融合策略常强迫异构模态无差别交互,可能引入冗余信息并破坏有价值的预训练表征。为解决该问题,本文从社会化学习视角重新审视多模态融合,提出适配DINOv3的A2DINOv3,这是一个具备社会化协作协议(SCP)的多专家协作框架。具体而言,RGB和红外分支被建模为异构专家,在独立保留各自专业知识的同时,通过选择性且受约束的交互交换互补信息。该设计减轻了有害的跨模态干扰,防止适配过程中预训练先验退化。此外,引入零初始化策略逐步激活跨模态协作,实现从模态特定学习到协作表征学习的平滑过渡。在四个多模态基准(包括航拍检测GAIIC、自动驾驶FLIR、低光照监控LLVIP及多样真实场景M3FD)上开展的大量实验表明,A2DINOv3在多模态目标检测中始终实现了超越现有技术的性能。
英文摘要
Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong representation capabilities, yet adapting them to multi-modal scenarios remains challenging. Existing dense cross-modal fusion strategies often force heterogeneous modalities to interact indiscriminately, which may introduce redundant information and disrupt the valuable pre-trained representations. To address this issue, we revisit multi-modal fusion from the perspective of socialized learning and propose adapter to DINOv3 (A2DINOv3), a multi-expert collaboration framework with a Socialized Collaboration Protocol (SCP). Specifically, RGB and infrared branches are modeled as heterogeneous experts that independently preserve their specialized knowledge while exchanging complementary information through selective and constrained interactions. This design mitigates harmful cross-modal interference and prevents degradation of pre-trained priors during adaptation. Furthermore, a zero-initialization strategy is introduced to gradually activate cross-modal collaboration, enabling a smooth transition from modality-specific learning to cooperative representation learning. Extensive experiments on four multi-modal benchmarks, including aerial detection (GAIIC), autonomous driving (FLIR), low-light surveillance (LLVIP), and diverse real-world scenarios (M3FD), demonstrate that A2DINOv3 consistently achieves state-of-the-art performance in multi-modal object detection.