发表机构
The Hong Kong Polytechnic University; Shandong University(香港理工大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对无监督动作分割的伪标签质量问题,提出SpecT-OT框架,结合SRP与TAR优化表示,在四个基准的15个指标中13个获最优结果,性能优于现有方法。
AI 中文摘要
无监督动作分割旨在在无动作标注的情况下发现潜在动作类别及其时间结构。基于最优传输(OT)的方法能提供帧到动作的结构化分配,但其伪标签质量根本上取决于构建传输代价所用的表示空间。我们认为,可靠的OT伪标签需要一种表示几何,该几何需同时对有判别性的动作变化敏感,且沿局部时间进程保持一致。基于此见解,我们提出SpecT-OT,一种基于不平衡最优传输伪标签概念的时空表示学习框架。SpecT-OT引入了频谱重参数化投影器(SRP),其用固定傅里叶基和可学习系数参数化投影器权重,以改进对快速变化的判别性特征的建模;还引入了时间亲和性正则化(TAR),其对成对帧亲和性施加距离感知的无标签约束,以稳定局部时间结构。这两个组件共同生成更具判别性和时间稳定性的传输代价,为迭代表示学习提供更可靠的伪标签。在四个基准上的实验表明,与最先进方法相比,该方法表现强劲;SpecT-OT在15个指标中的13个上取得最佳结果,包括在Breakfast和Desktop Assembly上分别比基线提升4.1个点的MoF和7.4个点的F1。
英文摘要
Unsupervised action segmentation aims to discover latent action categories and their temporal organization without action annotations. Optimal transport-based methods provide structured frame-to-action assignments, however, their pseudo-label quality is fundamentally conditioned on the representation space used to construct the transport cost. We argue that reliable OT pseudo-labeling requires a representation geometry that is simultaneously sensitive to discriminative action changes and coherent along local temporal progressions. Based on this insight, we propose SpecT-OT, a spectral-temporal representation learning framework built upon an unbalanced optimal transport pseudo-labeling concept. SpecT-OT introduces a Spectral Reparameterization Projector (SRP), which parameterizes projector weights with fixed Fourier bases and learnable coefficients to improve the modeling of rapidly varying discriminative features, and Temporal Affinity Regularization (TAR), which imposes distance-aware, label-free constraints on pairwise frame affinities to stabilize local temporal structure. The two components jointly produce more discriminative and temporally stable transport costs, yielding more reliable pseudo-labels for iterative representation learning. Experiments on four benchmarks demonstrate strong performance compared with state-of-the-art methods. SpecT-OT achieves the best results on 13 of 15 metrics, including 4.1-point MoF and 7.4-point F1 gains over the baseline on Breakfast and Desktop Assembly, respectively.