arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04623cs.CV

扩散模型中的视觉锚定:多模态零样本骨骼动作识别

Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition

Zehao Bao, Shujun Guo, Bruce X. B. Yu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对零样本骨骼动作识别中模态融合的问题,提出TDSM-MM模型,通过生成式分类范式结合视觉锚定,在NTU数据集上取得优异零样本识别性能,为零样本学习提供新方向。

中文摘要 AI 辅助

零样本骨骼动作识别(ZSAR)存在这样的问题:当未见过的动作具有相似的骨骼关节动态但在对象或场景上下文方面存在差异时,识别结果仍不明确。RGB提供了这些缺失的线索,但现有的多模态方法通常维护独立的骨骼和RGB评分分支并融合它们的输出。由于未使用未标记的测试数据进行适应或融合校准,固定的融合权重无法捕捉类别对依赖的模态可靠性,而自适应规则缺乏目标侧反馈来决定哪个分支应占主导地位。我们通过生成式分类范式绕过了这个权重选择问题,其中每个类别的评分依据是文本条件去噪器预测添加到骨骼特征中的噪声的准确程度。这种表述将逐渐被破坏的骨骼与固定条件分开,允许RGB和文本共同调节单个类别评分函数,而不是产生独立的评分。我们将这一思路实例化为用于骨骼-文本匹配的多模态三元组扩散模型(TDSM-MM),在文本条件去噪Transformer中添加了一个非扩散的RGB条件标记,该标记在骨骼数据重建过程中充当稳定的视觉锚定。我们提出的TDSM-MM已通过大量实验进行了消融研究,在四个NTU-60/120数据拆分中的三个上取得了最佳的归纳准确率,并且在NTU-120 96/24上超越了转导式最先进方法(即71.3%对比69.1%),且无需测试时适应,这表明基于扩散的方法可成为零样本学习的有前景方向。

英文摘要

Zero-shot Skeleton Action Recognition (ZSAR) remains ambiguous when unseen actions share similar skeleton joint dynamics but differ in objects or scene context. RGB provides these missing cues, yet existing multimodal methods typically maintain independent skeleton and RGB scoring branches and fuse their outputs. Without using unlabeled test data for adaptation or fusion calibration, a fixed fusion weight cannot capture class-pair-dependent modality reliability, while an adaptive rule lacks target-side feedback for deciding which branch should dominate. We bypass this weight-selection problem via the classify-by-generation paradigm, where each class is scored by how accurately a text-conditioned denoiser predicts the noise added to the skeleton feature. This formulation separates the progressively corrupted skeleton from fixed conditioning, allowing RGB and text to jointly condition a single class-scoring function rather than produce independent scores. We instantiate this idea as Multimodal Triplet Diffusion for Skeleton-Text Matching (TDSM-MM), augmenting a text-conditioned denoising Transformer with a non-diffused RGB condition token that serves as a stable visual anchor during skeleton data reconstruction. Our proposed TDSM-MM has been ablated via extensive experiments and achieved the best inductive accuracy on three of four NTU-60/120 splits and surpasses the transductive state-of-the-art on NTU-120 96/24 (i.e., 71.3% vs. 69.1%), without test-time adaptation, suggesting that diffusion-based methods can be a promising direction for zero-shot learning.

发表机构

  • The University of Hong Kong(香港大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑