发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对声呐鱼类追踪数据集定制了多模态大语言模型工具Molmo2Fish,通过交互式预测修正工作流开展实验,发现其在鱼类追踪和轨迹修正上性能良好,但自然语言引导的融入仍需提升。
AI 中文摘要
计算机视觉越来越多地用于自动化大型生态数据集的识别任务,但多目标追踪等更复杂的任务仍面临挑战。当研究人员尝试将视觉模型融入生态工作流时,诸多研究探索了如何通过人在回路过程让不完美的预测发挥作用。我们提出一种处理不完美追踪预测的新方法,即通过与多模态大语言模型对话的交互式预测修正工作流,并针对声呐鱼类追踪数据集定制该方法,作为概念验证的初始尝试。我们在有引导和无引导任务中研究了该工具Molmo2Fish的性能,涵盖修正自身预测的轨迹和外部轨迹。结果表明,Molmo2Fish在鱼类追踪和轨迹修正任务上实现了高性能,但在融入自然语言引导方面仍有很大提升空间。代码和数据可在此httpsURL公开获取。
英文摘要
Computer vision is increasingly used to automate recognition tasks in large ecological datasets, but more complex tasks such as multi-object tracking continue to pose challenges. As researchers seek to incorporate vision models in ecology workflows, various lines of research have explored how to make imperfect predictions useful through human-in-the-loop processes. We propose a new approach to working with imperfect tracking predictions through an interactive prediction correction workflow taking place as a conversation with a multimodal large language model, which we tailor to a sonar fish tracking dataset as an initial proof of concept. We investigate the performance of the tool, Molmo2Fish, across guided and unguided tasks, correcting its own predicted tracks and external tracks. We find that Molmo2Fish achieves high performance on fish tracking and track correction tasks, but there is still much room to improve on incorporating natural language guidance. The code and data are publicly available at https://github.com/tidalove/molmo2fish.
Comments29 pages, 6 figures, to be published in Third Workshop on Computer Vision for Ecology at ECCV 2026