arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Prompting-MammAlps:用于相机陷阱数据的细粒度文本到视频检索

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, Devis Tuia, Alexander Mathis

arXiv 2607.09876首次发表:更新:

发表机构

Ecole Polytechnique Fédérale de Lausanne (EPFL); Max Planck Institute of Animal Behavior; University of Konstanz(洛桑联邦理工学院; 马克斯·普朗克动物行为研究所; 康斯坦茨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对相机陷阱数据自动检索视频的难题,提出Prompting-MammAlps基准及细粒度可解释TVR方法,训练视觉变压器并结合大语言模型处理查询,在基准测试中取得较好成绩,优于零样本VLM。

AI 中文摘要

从大型相机陷阱数据集中自动检索视频具有挑战性。基于大型视频语言模型的文本到视频检索(TVR)方法有潜力通过文本查询检索感兴趣的事件,但当前方法缺乏时空理解且对生态数据泛化性不佳。本文引入Prompting-MammAlps首个相机陷阱TVR基准,并提出细粒度且可解释的TVR方法。训练视觉变压器进行时空动作定位并转换为结构化文本描述视频,由基于大语言模型的编码代理处理查询以解析结构化文本并检索视频,利用大语言模型使用自定义解析库函数降低幻觉风险并提高可解释性。该方法在基准测试中取得了34%的基于集合的F1分数,而最佳零样本VLM仅为18%且缺乏可解释性。

英文摘要

Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large video-language models (VLMs) have potential to retrieve events of interest by describing them with simple text queries. However, current methods often lack spatiotemporal understanding and do not generalize well to ecological data. In this work, we introduce Prompting-MammAlps, the first camera-trap TVR benchmark, and propose a fine-grained and interpretable TVR method. Specifically, we trained a vision transformer to perform spatiotemporal action localization, and convert its output to structured text, describing each video. Independently, ethology-inspired queries are processed by a Large-Language Model (LLM) based coding agent to parse the structured text per video and retrieve videos accordingly. We harnessed the LLM to use functions from a custom parsing library to minimize the risk of LLM hallucinations and to improve method interpretability. This retrieval approach applied on the Prompting-MammAlps benchmark achieved a set-based F1-score of 34\% on a test set of 135 ecologically-relevant queries and 775 candidate videos. In comparison the best zero-shot VLM achieved a F1-score of 18\%, while also lacking interpretability. Project page: https://cnai.epfl.ch/prompting-mammalps

CommentsAccepted at ECCV 2026; Project page: https://cnai.epfl.ch/prompting-mammalps

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑