arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LGFN:用于伪装目标检测的轻量级门控RGB-偏振融合与模态可用性条件化

LGFN: Lightweight Gated RGB-Polarization Fusion with Modality-Availability Conditioning for Camouflaged Object Detection

Zhuangfan Huang, Xiaosong Li, Yang Liu, Tao Ye, Haishu Tan

arXiv 2609.12798首次发表:更新:

发表机构

Foshan University; China University of Mining and Technology-Beijing(佛山大学; 中国矿业大学(北京))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LGFN轻量级门控RGB-偏振融合框架,通过模态路由和条件化门控处理伪装目标检测,在PCOD_1200上取得最优性能,并显著降低计算开销。

AI 中文摘要

伪装目标检测(COD)是智能光学感知中的一项重要工程任务,但当目标与周围环境高度相似时,该任务仍具挑战性。偏振成像提供了互补的物理线索,而现有方法通常假设固定的多模态输入配置,并将偏振内协调与红绿蓝(RGB)和偏振表示之间的交互纠缠在一起。我们提出LGFN,一种轻量级门控RGB-偏振融合框架,支持分别优化的仅RGB和偏振辅助配置。确定性模态路由器根据偏振可用性选择适当的配置。在多模态配置中,可用性条件化的模态门校准可用的偏振分支;门控偏振中心协调学习的线性偏振度(DoLP)和偏振角(AoP)表示与显式偏振线索;RGB-偏振交叉融合通过受控残差交互将协调表示引入RGB层级。多模态配置在推理时既不需要依赖样本的统计量,也不需要手工设计的质量描述符。在完整的230图像PCOD_1200测试集上,仅RGB配置实现了0.0090的平均绝对误差、0.8806的Dice分数和0.8144的交并比,在所评估的基于RGB的方法中,在所有六个指标上均取得最佳结果。在常见的局部重评估协议下,多模态配置在所有六个指标上优于PolarNet和IPNet。相对于IPNet,其参数数量、浮点运算次数和延迟分别减少了53.1%、73.6%和63.0%。

英文摘要

Camouflaged object detection (COD) is an important engineering task in intelligent optical perception, but it remains challenging when targets closely resemble their surroundings. Polarization imaging provides complementary physical cues, whereas existing methods typically assume fixed multimodal input configurations and entangle intra-polarization coordination with interaction between red-green-blue (RGB) and polarization representations. We propose LGFN, a lightweight gated RGB-polarization fusion framework supporting separately optimized RGB-only and polarization-assisted configurations. A deterministic Modality Router selects the appropriate configuration according to polarization availability. In the multimodal configuration, an availability-conditioned Modality Gate calibrates the available polarization branches; the Gated Polarization Hub coordinates learned degree of linear polarization (DoLP) and angle of polarization (AoP) representations with explicit polarization cues; and RGB-Polarization Cross Fusion introduces the coordinated representation into the RGB hierarchy through controlled residual interaction. The multimodal configuration requires neither sample-dependent statistics nor handcrafted quality descriptors during inference. On the complete 230-image PCOD_1200 test set, the RGB-only configuration achieves a mean absolute error of 0.0090, a Dice score of 0.8806, and an intersection over union of 0.8144, obtaining the best results on all six metrics among the evaluated RGB-based methods. Under a common local reevaluation protocol, the multimodal configuration outperforms PolarNet and IPNet on all six metrics. Relative to IPNet, it reduces the parameter count, floating-point operations, and latency by 53.1%, 73.6%, and 63.0%, respectively. The source code will be available at https://github.com/1hzf/LGFN.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑