arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15628eess.IV

MSTF-Net:一种通过模态稳健、尺度自适应和一致融合的面向无人机的多光谱视频分割方法

MSTF-Net: A UAV-Oriented Multi-Spectral Video Segmentation Method via Modality-Robust, Scale-Adaptive, and Consistent Fusion

Chenwei Wang, Zhida Wang, Zelin Li, Elif Ozden-Yenigun, Houde Liu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对无人机多光谱视频分割面临的模态融合困境和时间变化问题,提出MSTF-Net框架,通过模态稳健、尺度自适应融合有效建模跨模态融合与时间一致性,在公共数据集实验中取得出色分割性能。

中文摘要 AI 辅助

多光谱视频分割对于无人机应用中的稳健场景理解至关重要,如城市规划、土地利用监测等。RGB和热模态融合能提供互补信息,但存在模态融合困境和时间变化两大挑战。本研究提出MSTF-Net,一个用于多光谱视频分割的模态稳健尺度自适应融合框架。其模态空间互补抑制与增强模块通过跨模态注意力生成统一实例查询,多尺度时间跨模态语义一致性模块基于帧距离自适应调整时间感受野。在公共RGB-T数据集上的消融实验表明,MSTF-Net实现了最先进的分割性能,在小目标、遮挡和模态退化等具有挑战性的条件下表现出色,在MVSeg数据集上达到56.42%的平均交并比,在CART数据集上达到51.80%的平均交并比。

英文摘要

Multi-spectral video segmentation is essential for robust scene understanding in unmanned aerial vehicle (UAV) applications such as city planning, land use monitoring, traffic monitoring, and crowd estimation. While the fusion of RGB and thermal modalities offers complementary information for perception under varying lighting and visibility conditions, two fundamental challenges remain: (1) the modal fusion dilemma, arising from significant discrepancies between RGB and thermal features that obscure complementary cues, and (2) temporal variation, induced by rapid motion and viewpoint changes on UAV platforms, which leads to appearance inconsistency and misalignment across frames. To address these issues, this study proposed MSTF-Net, a modality-robust scale-adaptive fusion framework for multi-spectral video segmentation that effectively models cross-modal fusion and temporal consistency. The Modality Spatial Complementary Suppression and Enhancement (MSCSE) module generates unified instance queries via cross-modal attention and suppresses modality-specific noise using residual-guided discrepancy filtering and consistency constraints. To model temporal dynamics, the Multi-scale Temporal Cross-modality Semantic Consistency (MTCSC) module adaptively adjusts the temporal receptive field based on frame distance, capturing both coarse global context and fine local structure across time. Extensive ablation experiments on public RGB-T datasets demonstrate that MSTF-Net achieves state-of-the-art segmentation performance, especially under challenging conditions such as small targets, occlusion, and modality degradation. Specifically, reached 56.42\% mIoU on the MVSeg dataset and 51.80\% mIoU on the CART dataset, respectively.

↑