arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于正未标记海洋物种检测的具有提议重排和分数融合的解耦管道

Decoupled Pipeline with Proposal Reranking and Score Fusion for Positive-Unlabeled Marine Species Detection

Robert James Brock, Sebastian Maximilian Krupa, Jason Kahei Tam

arXiv 2607.18700首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FathomNetCLEF 2026竞赛面临训练标签稀疏等问题,DS@GT ARC团队开发多阶段系统,用冻结检测器生成提议,经推理、分类和分数融合排序,在102个团队中排第12,实验表明特定策略比微调检测器等更有效。

AI 中文摘要

FathomNetCLEF 2026竞赛在正未标记评估设置下结合了水下目标检测和细粒度海洋物种分类。训练标签稀疏,隐藏测试集与训练图像分布不同,带来标注不完整和源转移挑战。我们描述了DS@GT ARC为此设置开发的多阶段系统,最终模型使用冻结的Megalodon YOLOv8x检测器生成提议,结合全局和分块推理及边缘过滤,用LoRA微调的DINOv3 ViT-H分类器对提议作物分类,通过加权几何融合对预测排序,该系统在102个团队中排第12。相关变体添加局部训练的有效性头部,提升了公开排行榜和代理评估性能,但降低了私有排行榜性能。实验表明训练衍生验证和仅检测器指标不可靠,应使用代理数据集验证和比较,并结合排行榜反馈和针对性消融。保留提议召回率、避免过度过滤和改进下游排序比微调检测器或直接在有噪声伪标签上训练更有效。

英文摘要

The FathomNetCLEF 2026 competition combines underwater object detection and fine-grained marine species classification under a positive-unlabeled evaluation setting. The provided training labels are sparse, while the hidden test set is out-of-distribution relative to the training imagery, creating both annotation incompleteness and source-shift challenges. We describe DS@GT ARC's multi-stage system developed for this setting while keeping model training restricted to the data provided by the competition. The final private-leaderboard model uses a frozen Megalodon YOLOv8x detector as a class-agnostic proposal generator, combines global and tiled inference with tile-edge filtering, classifies expanded proposal crops with a LoRA-finetuned DINOv3 ViT-H classifier, and ranks predictions using weighted geometric fusion of detector and classifier confidence. This system placed 12th out of 102 teams. A closely related variant added a locally trained TTN-inspired validity head as a light reranking signal, improving public-leaderboard and proxy-evaluation performance but slightly reducing private-leaderboard performance. Across experiments, the strongest lesson was that train-derived validation and detector-only metrics were not reliable enough for model selection. Instead, we used proxy datasets only for validation and comparison, and combined those signals with leaderboard feedback and targeted ablations. These experiments showed that reserving proposal recall, avoiding over-aggressive filtering, and improving downstream ranking were more effective than fine-tuning the detector or directly training on noisy pseudo-labels. Code: https://github.com/dsgt-arc/fathomnetclef-2026.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑