arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25500cs.CV

mbariml:一个将深海图像和视频转化为目标检测训练数据的策展流水线

mbariml: a curation pipeline for turning deep-sea imagery and video into object-detection training data

Lonny Lundsten, Kevin Barnard, Dave Caress

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出mbariml,一个基于Python的深海图像和视频目标检测训练数据策展流水线,通过YOLO检测、按视觉相似性分组批量审查及代表性帧选择,提升数据标注效率与模型性能。

中文摘要 AI 辅助

训练数据的数量和质量极大地影响目标检测模型的性能,无论模型架构如何。当在深海视频和图像上使用目标检测模型时,其中感兴趣的目标(主要是生物)稀疏、微弱且难以识别,目标检测器性能的增量改进可能需要一种迭代式的数据标注和管理方法。本文介绍了mbariml,一个基于Python的视频和图像分析流水线,围绕数据标注管理过程构建。mbariml使用Ultralytics YOLO检测模型,在静态图像或视频上运行该模型,将每次检测存储为可审查的关注区域,按视觉相似性对这些区域进行分组,以便人类可以批量接受或拒绝它们,并将结果导出为训练数据、统计数据、图像侧车文件和其他元数据。人工审查阶段是设计的核心:标注者可以验证、重新标注、调整大小、删除和绘制全新的定位框,并且每一次编辑都会写回检测器写入的同一个数据库。视频受到特别关注:该软件将每个跟踪器产生的轨迹视为临时观察,并选择一个代表性帧,而不是保留轨迹中的每次检测。我们逐阶段描述该流水线,包括用于轨迹观察选择的操作性中间三分之一启发式方法。

英文摘要

Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. For object detection in deep-sea video and imagery, where the objects of interest (primarily organisms) are sparse, faint, and hard to identify, incremental improvements to detector performance may require an iterative approach to data labeling and management. This paper presents mbariml, a Python-based video and image analysis pipeline built around the data labeling and management process. mbariml uses an Ultralytics YOLO detection model, runs it over still images or video, stores every detection as a reviewable region of interest, groups those regions by visual similarity so that a human can accept or reject them in bulk, and exports the result as training data, statistics, image sidecars, and additional metadata. Existing YOLO and Pascal VOC datasets can be imported into the same database, so a legacy training set can be reviewed, extended, and re-exported alongside new detections. The human review stage is the center of the design: an annotator can validate, relabel, resize, delete, and draw entirely new localizations, optionally assisted by the SAM3 segmentation model, and every one of those edits is written back to the same database the detector wrote to. Video receives particular attention: the software treats each tracker-produced track as a provisional observation and selects one representative frame instead of retaining every detection in the track. We describe the pipeline stage by stage, including how each track's representative frame is chosen (from a user-selected third of the track, the middle by default), which we examine on 684 tracks from seafloor video.

发表机构

  • Monterey Bay Aquarium Research Institute(蒙特雷湾水族馆研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑