arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FIS-OT:用于无监督动作分割的特征诱导最优传输

FIS-OT: Feature-Induced Optimal Transport for Unsupervised Action Segmentation

Linxiang Peng, Xinyao Qin, Jinhan Li, Di Yang, Jiangtao Wang

arXiv 2608.29980首次发表:更新:

发表机构

Suzhou Institute for Advanced Research, University of Science and Technology of China; School of Artificial Intelligence and Data Science, University of Science and Technology of China; Suzhou Big Data & AI Research and Engineering Center; Shanghai Key Laboratory of Data Science(中国科学技术大学苏州高等研究院; 中国科学技术大学人工智能与数据科学学院; 苏州大数据与人工智能研究工程中心; 上海数据科学重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有最优传输方法忽略局部信息、易受伪标签噪声影响的问题,本文提出FIS-OT框架,通过FEG模块、特征诱导残差结构先验及循环优化环,在三个数据集上验证了方法的有效性。

AI 中文摘要

无监督动作分割是一项具有挑战性的任务,需要在无标签的视频中找到动作类别与边界。现有最优传输(OT)方法采用全局约束,导致其忽略局部信息的利用;此外,现有最优传输架构易出现确认偏差,因为它们过度信任自身生成的伪标签,使得模型在训练早期就从噪声中学习。为解决这些问题,本文提出FIS-OT,这是一种新型特征诱导结构化最优传输框架。首先,引入特征增强生成器(FEG)模块,作为内部正则化器,通过三元组损失捕获局部一致性,提供不依赖噪声伪标签的鲁棒监督;其次,提出特征诱导残差结构先验,结合固定时间骨干网络与动态特征相似度,该设计既保证时间连续性,又使求解器能适配复杂动作结构;最后,建立循环优化循环,使局部特征学习与全局结构对齐相匹配。在三个数据集上开展的大量实验验证了本文方法的有效性。

英文摘要

Unsupervised action segmentation is a challenging task. It involves finding action categories and boundaries in videos without labels. Existing Optimal Transport (OT) methods use global constraints. This causes them to overlook the use of local information. Furthermore, existing Optimal transport architectures are prone to confirmation bias because they overly trust the pseudo-labels they generate. This causes models to learn from noise in the early training stages. To address these issues, we propose FIS-OT. It is a novel Feature-Induced Structured Optimal Transport framework. First, we introduce a Feature Enhanced Generator (FEG) module. It serves as an internal regularizer. By using triplet loss, FEG captures local consistency. It provides robust supervision that is independent of noisy pseudo-labels. Second, we propose a Feature-Induced Residual Structural Prior. This combines a fixed temporal backbone with dynamic feature similarities. This design ensures temporal continuity. It also allows the solver to adapt to complex action structures. Finally, we establish a cyclic optimization loop. This aligns local feature learning with global structural alignment. Extensive experiments on the three datasets show the effectiveness of our method.

CommentsAccepted by ICME2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑