arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ConsensusTAS:面向长时程施工视频的自监督时序动作分割

ConsensusTAS: Self-Supervised Temporal Action Segmentation for Long-Horizon Construction Videos

Xiaoshan Zhou, Yafei Sun

arXiv 2608.24043首次发表:更新:

发表机构

School of Project Management, University of Sydney; School of Civil and Environmental Engineering, University of New South Wales(悉尼大学项目管理学院; 新南威尔士大学土木与环境工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出自监督时序动作分割模型ConsensusTAS,无需标注且可在CPU运行,在公开数据集及真实施工视频上的性能优于现有最优方法,可用于人机协同与视频监控。

AI 中文摘要

识别时序施工活动对人机协同工作至关重要;例如,机器人能够理解工人当前及即将开展的动作,并及时提供工具递送或物理支持。然而,尽管施工工人活动识别领域已有大量研究,现有研究仅局限于对攀爬、搬运、行走等活动类别进行分类,而非从长时程序列中识别细粒度的活动过渡。解决该问题颇具挑战性,因为标注长施工视频中的动作时间边界极为耗时。本研究提出ConsensusTAS,这是一种无标签的自监督学习方法,通过利用候选分割的内部共识,将连续视频流分割为不同的活动阶段。我们在三个公开数据集上对该算法进行了评估,其性能优于现有最优方法:在GTEA数据集上取得F1@10为73.08,在Breakfast数据集上取得F1@10为64.33,在Assembly101的静态相机视频上取得F1@50为33.50。我们还在真实施工视频上进行了测试,事后评估显示,该模型成功识别并分割了砌砖复合活动内的动作,如在砖块上涂抹砂浆、放置砖块、按压以及对齐。与需要计算密集型大视觉语言模型的其他时序动作分割模型相比,我们的方法可在CPU上运行,为移动机器人平台上的视频监控和人机协同提供了实用价值。

英文摘要

Recognizing sequential construction activities is important for collaborative human-robot work; for example, robots are able to understand workers' current and upcoming actions and provide timely tool delivery or physical support. However, despite extensive research on construction worker activity recognition, existing studies have been limited to classifying activity categories, such as climbing, lifting, and walking, instead of recognizing fine-grained activity transitions from long-horizon sequences. Addressing this problem is challenging because annotating action temporal boundaries in long construction videos is time-consuming. In this study, we propose ConsensusTAS, a label-free, self-supervised learning approach to segment continuous video streams into distinct activity phases by exploiting the internal consensus of candidate segmentations. We evaluated our algorithm on three public datasets, where it outperformed state-of-the-art methods, achieving an F1@10 of 73.08 on GTEA, an F1@10 of 64.33 on Breakfast, and an F1@50 of 33.50 on static-camera videos from Assembly101. We also tested it on real-world construction videos, where post-hoc evaluation showed that the model successfully recognized and segmented actions within the composite activity of bricklaying, such as spreading mortar on a brick, placing the brick, pressing, and aligning. Compared with other temporal action segmentation models that require computationally intensive large vision-language models, our method can run on a CPU, which provides practical value for video surveillance and human-robot collaboration on mobile robotic platforms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑