arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyTBE: 用于跨域红外小目标检测的双曲目标-背景专家模型

HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection

Aohua Li, Jin Kuang, Yubing Lu, Pingping Liu

arXiv 2608.05771首次发表:更新:

AI 中文总结

针对跨域红外小目标检测中目标-背景关系偏移导致的性能下降问题,提出HyTBE模型,通过关系干预、双曲关系建模及MoE适配器实现更强的跨域泛化。

AI 中文摘要

红外小目标检测(IRSTD)在域一致性评估下已取得显著进展,但当推广至未见过的红外域时,检测器性能往往会明显下降。现有方法主要通过增强目标响应和抑制背景干扰来提升检测效果。然而,当仅在有限的源域集合上训练时,其学习到的决策规则必然建立在有限范围的源域目标-背景关系模式上。我们将这种跨域失败表述为目标-背景关系偏移:未见过的域可能呈现训练过程中未观察到的关系模式,从而削弱从源域学习到的判别能力。为解决该问题,我们提出HyTBE,一种双曲目标-背景专家模型,该模型可扩展源域关系模式并利用显式关系线索自适应调整视觉表示。目标-背景关系干预会选择性地扰动目标或背景,在保持有效监督的同时拓宽训练期间可观察到的关系模式。随后,双曲关系建模将多尺度视觉线索映射到庞加莱球,并根据每个特征标记与目标和背景锚点的相对距离来表征其目标-背景关系。双曲门控MoE适配器进一步利用这些双曲关系表示来校准多尺度视觉特征,并聚合针对不同关系模式的专家特定特征修正。在NUAA-SIRST、NUDT-SIRST和IRSTD-1K上进行的留一域实验表明,HyTBE比竞争性基线实现了更强的跨域泛化能力。

英文摘要

Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains. Existing methods primarily improve detection by enhancing target responses and suppressing background interference. However, when trained on only a limited set of source domains, their learned decision rules are inevitably established from a restricted range of source-domain target-background relation patterns. We formulate this cross-domain failure as target-background relation shift: unseen domains may exhibit relation patterns that are not observed during training, thereby weakening the discriminative capability learned from the source domains. To address this problem, we propose HyTBE, a Hyperbolic Target-Background Expert model that expands source-domain relation patterns and adaptively adjusts visual representations using explicit relation cues. The Target-Background Relation Intervention selectively perturbs either targets or backgrounds, broadening the observable relation patterns during training while maintaining valid supervision. Subsequently, the Hyperbolic Relation Modeling maps multi-scale visual cues into a Poincaré ball and characterizes the target-background relation of each feature token according to its relative distances to the target and background anchors. The Hyperbolic-guided MoE Adapter further uses these hyperbolic relation representations to calibrate multi-scale visual features and aggregate expert-specific feature corrections for different relation patterns. Leave-one-domain-out experiments on NUAA-SIRST, NUDT-SIRST, and IRSTD-1K demonstrate that HyTBE achieves stronger cross-domain generalization than competitive baselines.

Comments15 pages, 9 figures, 9 tables. Code: https://github.com/PepperCS/HyTBE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑