arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究视频Transformer中的时空冗余性以实现碰撞预判

Investigating Spatiotemporal Redundancy in Video Transformer for Collision Anticipation

Xiaoshan Zhou

arXiv 2610.04727首次发表:更新:

发表机构

University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析视频Transformer在碰撞预判中的时空冗余,提出时间令牌合并与神经元剪枝方法,实现1.76倍加速且性能几乎无损。

AI 中文摘要

在工人与设备接近监测中,视频Transformer被广泛用于碰撞预判,并展现出强大的性能。然而,其准确性伴随着巨大的计算需求,这与移动机器人上低延迟推理的需求以及建筑领域追求低碳计算的目标形成了矛盾。为解决这一问题,本研究探究了既有视频Transformer内部计算冗余的位置,以及在不显著降低预测性能的前提下,这些冗余是否可以被移除。使用VideoMAEv2-Base模型在Nexar碰撞预测数据集上,我们首先考察了碰撞相关信息如何随网络深度演变,然后研究了两种互补的冗余形式:多层感知器(MLPs)中的结构化容量冗余和令牌流中的时空冗余。线性探针显示,可解释的运动线索,包括光流幅度、逼近和后退,在中间层最易获取,而碰撞标签判别力在最终层增强。令牌冗余具有轴特异性:在较深层,相邻时间令牌的相似度达到0.970,而空间相似度降至0.297,表明时间维度上的冗余远大于空间维度。利用这一不对称性,时间令牌合并将骨干网络计算量从356.99 GFLOPs降至178.50 GFLOPs,每个片段的延迟从12.69毫秒降至7.21毫秒,实现了1.76倍的加速,而平均精度均值仅从0.7478变为0.7443。基于重要性保留50%的MLP单元,AUC保持为0.753,而在匹配的随机保留下AUC为0.529,并揭示了剪枝在判别排名崩溃之前改变了分数校准。这些发现为通过有针对性的时间令牌压缩和神经元剪枝追求更快速算法开辟了新途径。

英文摘要

In worker-equipment proximity monitoring, video transformers are widely used for collision anticipation and have demonstrated strong performance. However, their accuracy comes with substantial computational demands, creating a tension with the need for low-latency inference on mobile robots and the pursuit of lower-carbon computation in construction. To address this, this study investigates where computation within an established video transformer is redundant and whether that redundancy can be removed without materially degrading predictive performance. Using VideoMAEv2-Base on the Nexar Collision Prediction dataset, we first examine how collision-relevant information evolves across network depth and then investigate two complementary forms of redundancy: structured capacity redundancy in multilayer perceptrons (MLPs) and spatiotemporal redundancy in the token stream. Linear probes show that interpretable motion cues, including flow magnitude, looming, and approach versus retreat, are most accessible at intermediate layers, whereas collision-label discrimination strengthens toward the final layer. Token redundancy is axis-specific: adjacent temporal-token similarity reaches 0.970 in later layers, while spatial similarity falls to 0.297, indicating substantially greater redundancy across time than across space. Exploiting this asymmetry, temporal token merging reduces backbone computation from 356.99 to 178.50 GFLOPs and latency from 12.69 to 7.21 ms per clip, a 1.76x speedup, while mean average precision changes only from 0.7478 to 0.7443. Importance-guided retention of 50% of MLP units preserves an AUC of 0.753, compared with 0.529 under matched random retention, and reveals that pruning alters score calibration before discriminative ranking collapses. These findings establish a new pathway for pursuing faster algorithms through targeted temporal token compression and neuron pruning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑