发表机构
Research Institute for Science & Technology, Tokyo University of Science(东京理科大学科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对无人机视频运动模糊问题,提出自适应潜在尺度选择器与多帧对齐可学习门控模块,实现高效去模糊,提升无人机目标检测性能,适用于鲁棒航拍视觉任务。
AI 中文摘要
无人机(UAV)在从灾害应对到交通监控的各类场景中发挥着关键作用。然而,航拍视频画面常因快速飞行机动、振动及相机平移出现严重运动模糊,这会显著降低目标检测等下游任务的性能。我们的目标是探索一种计算高效且有效的视频去模糊方法,以提升无人机目标检测性能。为降低计算成本,我们首先提出自适应潜在尺度选择器(Adaptive Latent Scale Selector),其可根据无人机运动强度动态调整潜在空间分辨率,从而在细节保留与推理效率间实现平衡。为确保时间一致性,我们引入多帧对齐与可学习门控模块(Multi-Frame Alignment and Learnable Gating),用于对前序帧进行配准与门控,使模型仅融合相关的时间信息,同时抑制未对齐或无信息的特征。我们的方法可有效从无人机视频流中恢复清晰细节。在真实无人机基准数据集上开展的大量实验表明,我们的方法不仅实现了更优的去模糊性能,还显著提升了目标检测准确率,使其高度适用于鲁棒航拍视觉任务。
英文摘要
Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to traffic surveillance. However, aerial video footage often suffers from severe motion blur due to rapid flight maneuvers, vibrations, and camera panning, which can significantly degrade downstream tasks such as target detection. Our goal is to explore a computationally-efficient and effective video deblurring approach to enhance UAV target detection performance. To reduce computational cost, we first propose an Adaptive Latent Scale Selector that dynamically adjusts the latent space resolution according to the intensity of UAV motion, thus balancing detail preservation with inference efficiency. To ensure temporal consistency, we introduce a Multi-Frame Alignment and Learnable Gating module to warp and gate the preceding frames, allowing the model to fuse only relevant temporal information and suppress misaligned or uninformative features. Our method can effectively recover sharp details from the UAV video stream. Extensive experiments on real UAV benchmarks demonstrate that our method not only yields superior deblurring performance but also significantly boosts target detection accuracy, making it highly applicable to robust aerial vision tasks.
Comments8 pages, 8 figures. Published in the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)