arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

帧到全景定位与上下文感知采样用于智能游艇码头测试平台中的场景特定船舶检测

Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

Ignat Romanov, Andreas Hadjipieris, Neofytos Dimitriou

arXiv 2609.29447首次发表:更新:

发表机构

University of Nicosia; Cyprus Marine and Maritime Institute(尼科西亚大学; 塞浦路斯海洋与海事研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种帧到全景定位与上下文感知采样流水线,将历史PTZ海事视频缩减为紧凑训练集,用于场景特定船舶检测,在智能游艇码头测试平台中实现99.5%数据缩减和94.78%的AP50。

AI 中文摘要

智能海事基础设施提供对异构传感流的持续访问,支持重复实验、数字孪生开发和基于AI的海事服务。然而,仅靠传感硬件不足以进行场景特定的模型开发:历史视频流还必须进行空间索引、上下文化,并缩减为用于注释的信息子集。本文提出了一种帧到全景定位和上下文感知采样流水线,用于缺乏可靠平移、倾斜和缩放元数据的历史PTZ海事视频中的船舶检测。主要贡献是一种端到端的数据策展方法,从历史PTZ视频中恢复相机视角信息,并将其与环境上下文和视觉多样性相结合,构建紧凑的场景特定训练集。具体而言,帧使用SuperPoint和LightGlue在参考全景图上进行定位,丰富天气和太阳状态元数据,并通过多样性采样选择以保留相机视角和环境条件下的变化。第二阶段上下文感知阶段针对地平线附近代表性不足的远距离船舶案例,使用瓦片级视觉嵌入和高斯混合模型聚类。在CMMI MDigi-I智能游艇码头测试平台中应用,所提出的流水线将40,718个候选帧减少到220张用于注释的图像,对应99.5%的减少。在此子集上微调的YOLO26-m检测器在序列分组五折交叉验证下实现了平均AP50为94.78%±0.51%和平均AP50-95为75.10%±1.73%。这些结果表明,高度冗余的基础设施视频流可以转化为紧凑、空间和上下文多样的训练集,用于场景特定检测器适应,同时大幅减少注释工作。

英文摘要

Smart maritime infrastructures provide continuous access to heterogeneous sensing streams, enabling repeated experimentation, digital-twin development, and AI-based maritime services. However, sensing hardware alone is not sufficient for scene-specific model development: historical video streams must also be spatially indexed, contextualized, and reduced to informative subsets for annotation. This paper presents a frame-to-panorama localization and context-aware sampling pipeline for ship detection in historical PTZ maritime video lacking reliable pan, tilt, and zoom metadata. The main contribution is an end-to-end data-curation approach that recovers camera-view information from historical PTZ video and combines it with environmental context and visual diversity to construct compact, scene-specific training sets. Specifically, frames are localized on a reference panorama using SuperPoint and LightGlue, enriched with weather and solar-state metadata, and selected through diversity sampling to preserve variation across camera view and environmental conditions. A second context-aware stage targets under-represented distant-vessel cases near the horizon using tile-level visual embeddings and Gaussian Mixture Model clustering. Applied within the CMMI MDigi-I Smart Marina testbed, the proposed pipeline reduces 40,718 candidate frames to 220 images for annotation, corresponding to a 99.5% reduction. A YOLO26-m detector fine-tuned on this subset achieves a mean AP50 of 94.78% $\pm$ 0.51% and a mean AP50-95 of 75.10% $\pm$ 1.73% under sequence-grouped five-fold cross-validation. These results demonstrate that highly redundant infrastructure video streams can be transformed into compact, spatially and contextually diverse training sets for scene-specific detector adaptation while substantially reducing annotation effort.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑