arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

D-CLOT:用于无监督动作分割的双闭环最优传输

D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation

Elena Bueno-Benito, Mariella Dimiccoli

arXiv 2608.05877首次发表:更新:

发表机构

Institut de Robòtica i Informàtica Industrial, CSIC-UPC(机器人学与工业信息学研究所(西班牙国家研究委员会-加泰罗尼亚理工大学联合机构))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无监督动作分割中最优传输方法的表示-原型不一致瓶颈,提出D-CLOT及其变体D-CLOT_B,通过图约束模块和动作原型更新机制,在多个基准上取得显著性能提升,并建立Assembly101首个无监督基线。

AI 中文摘要

最优传输(OT)已成为无监督动作分割的有效框架。然而,在现有的基于OT的方法中,定义OT代价的潜在动作原型并未从优化后的帧几何结构中重新估计,而是仅通过伪标签损失的梯度演化。我们将这种“表示-原型不一致”确定为核心瓶颈,尤其在模糊过渡区域以及短动作或不频繁动作中更为明显。为解决该问题,我们基于最新提出的CLOT构建,CLOT会基于估计的片段嵌入优化帧嵌入,我们进一步从优化后的帧嵌入中重新估计动作原型。具体而言,我们引入图约束模块,通过保留编码器输出的局部邻域几何结构来正则化OT优化后的帧和片段表示。随后的动作嵌入细化步骤会定期将原型重新锚定到该稳定的表示几何结构上。我们研究了两种实例,它们共享相同的骨干网络、图模块和目标:D-CLOT使用k-means更新原型,而D-CLOT_B则根据优化后的传输计划将原型更新为OT重心,从而得到与当前传输几何结构一致的感知分配的原型更新。在五个已建立的基准上,两种变体均比CLOT提升了片段级质量,每视频的F1得分最高提升12.7,mIoU(YTI)最高提升10.2;活动级的F1得分最高提升8.9(FS-Eval)。我们进一步在Assembly101上建立了首个无监督动作分割基线,Assembly101是一个比以往工作中常用基准更精细的程序型基准。大量的消融实验和敏感性分析表明,两种细化机制具有互补性且表现稳健。

英文摘要

Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation. Yet, in existing OT-based methods, the latent action prototypes that define the OT costs are not re-estimated from the refined frame geometry. Instead, they evolve solely through gradients from the pseudo-label loss. We identify this \emph{representation--prototype inconsistency} as a central bottleneck, particularly around ambiguous transitions and for short or infrequent actions. To address this issue, we build on the recently introduced CLOT, which refines frame embeddings based on estimated segment embeddings, and further re-estimates the action prototypes from the refined frame embeddings. Specifically, we introduce a graph-constrained module that regularizes the OT-refined frame and segment representations by preserving the local neighborhood geometry of the encoder output. An action-embedding refinement step then periodically re-anchors the prototypes to this stabilized representation geometry. We study two instantiations that share the same backbone, graph module, and objective: D-CLOT updates the prototypes using $k$-means, whereas D-CLOT$_{B}$ updates them as OT barycenters weighted by the refined transport plan, yielding an assignment-aware prototype update consistent with the current transport geometry. Across five established benchmarks, both variants improve segment-level quality over CLOT, with per-video gains of up to $+12.7$ F1 and $+10.2$ mIoU (YTI) and activity-level gains of up to $+8.9$ F1 (FS-Eval). We further establish the first unsupervised action-segmentation baseline on Assembly101, a procedural and substantially more fine-grained benchmark than those commonly used in prior work. Extensive ablations and sensitivity analyses demonstrate that the two refinement mechanisms are complementary and robust.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑