AI 中文总结
该研究提出一种两阶段开源工作流,基于Echopype和Echoregions库构建可扩展的机器学习就绪回声测深数据集,含数据划分与掩码创建,提供教程验证其可行性。
AI 中文摘要
回声测深仪(即高频主动声呐系统)已成为渔业或生态调查中量化和测绘海洋生物分布的标准工具。传统回声测深数据分析通常依赖人工标注声呐图,声呐图是由回波强度形成的声呐图像。过去十年,随着回声测深数据量呈指数级增长,以声呐图为图像的机器学习(ML)方法也相应发展。然而,声呐图并非简单图像:它们关联着特定的时空坐标,这对与调查事件、人工标注及其他海洋学数据集的对齐至关重要。我们提出了一种可推广的两阶段工作流,用于构建适用于基于断面调查的声呐图的机器学习分析就绪数据集:(1)声学数据按断面划分;(2)根据用户定义的统一时空回波数据网格,从标注中创建掩码。重要的是,地理空间坐标和海洋学测量等辅助信息会在处理阶段间传递,以保留下游分析所需的关键上下文信息。我们基于两个开源软件库Echopype和Echoregions,使用两个渔业调查示例数据集验证了工作流的可扩展性。此外,我们还提供了可执行教程,指导读者完成该工作流的计算实现。这些要素共同构成了一个可扩展且可推广的框架,用于创建面向机器学习应用的分析就绪回声测深数据集。
英文摘要
Echosounders, or high-frequency active sonar systems, have become standard tools for quantifying and mapping the distribution of marine organisms in fisheries or ecological surveys. Conventional echosounder data analysis often relies on human annotation of echograms, which are sonar imagery formed by echo intensity. Over the past decade, in parallel with the exponentially growing volume of echosounder data, there has been a corresponding increase in the development of machine learning (ML) methods that operate primarily on echograms as images. However, echograms are not simply images: they are associated with specific spatiotemporal coordinates that are essential for alignment with survey events, human annotations, and other oceanographic datasets. We present a generalizable two-stage workflow for constructing analysis-ready datasets for ML development tailored for echograms from transect-based surveys, in which (1) acoustic data are partitioned according to transect designation, and (2) masks are created from annotations referencing user-defined uniform spatiotemporal echo data grid. Importantly, ancillary information, such as geospatial coordinates and oceanographic measurements, is propagated across processing stages to preserve the essential contextual information for downstream analyses. We demonstrate the scalability of our workflow implementation based on two open-source software libraries, Echopype and Echoregions, using two example fisheries survey datasets. We additionally provide an executable tutorial that guides readers through the computational implementation of this workflow. Together, these elements provide a scalable and generalizable framework for creating analysis-ready echosounder datasets for ML applications.