arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于VLM描述比较的异常帧检测,用于提取特定专家操作和具有视频内自相似性的上下文决策场景

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

Ryo Sakai, Kaname Yokoyama

arXiv 2607.11957首次发表:更新:

AI 中文总结

本文针对关键基础设施维护中专家知识传承问题,提出基于VLM描述比较检测异常帧的方法,可提取特定专家操作及上下文决策场景,在模拟实验中该方法的提取率高于传统方法,有效发现含专家知识的候选场景。

AI 中文摘要

关键基础设施(如铁路和发电厂)的维护对于确保运行安全和可靠性至关重要。然而,熟练维护工人数量的减少凸显了将专家知识传授给经验不足工人的必要性。以往研究主要关注可观察动作的差异,而专家知识不仅体现在动作中,还体现在任务执行过程中的上下文决策中。本文提出一种检测两个任务视频之间异常帧的方法,以自动提取包含特定专家操作和上下文决策场景的候选场景。该方法使用视觉语言模型(VLM)生成逐帧视觉描述,基于两个视频描述比较计算的帧相似度提取特定专家操作,利用描述的视频内自相似性导出的片段相似度提取上下文决策场景。在涉及27个任务场景的模拟配电板维护实验中,该方法的动作候选提取率为65%,决策场景候选提取率为61%,优于传统方法(分别为59%和33%)。这些结果证明了该方法在发现包含专家知识的候选场景方面的有效性。

英文摘要

Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the declining number of skilled maintenance workers highlights the need to transfer expert know-how to less experienced workers. Previous studies have attempted to extract candidates of expert knowledge by comparing videos of manual-based work with those of expert workers, mainly focusing on differences in observable actions. However, expert know-how is often embedded not only in actions but also in contextual decision-making during task execution. This paper proposes a method that detects anomalous frames between two task videos to automatically extract candidate scenes containing expert-specific actions and contextual decision-making scenes. The method generates frame-wise visual descriptions using a vision-language model (VLM). Expert-specific actions are extracted based on frame similarities computed from description comparisons between two videos, while contextual decision-making scenes are extracted using segment similarities derived from intra-video self-similarity of the descriptions. In simulated distribution board maintenance experiments involving 27 task scenarios, the proposed method achieved extraction rates of 65% for action candidates and 61% for decision-scene candidates, improving over conventional methods that achieved 59% and 33%, respectively. These results demonstrate the effectiveness of the proposed approach in discovering candidate scenes containing expert know-how.

Comments17 pages, 11 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑