arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GMoT:用于基于多模态大语言模型的细粒度微手势视频推理的门控运动感知令牌化

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

Taorui Wang, Wei Xia, Hui Ma, Zijia Song, Jiayu Zhang, Zeheng Wang, Yong Xu, Zitong Yu

arXiv 2607.16322首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Great Bay University; Dongguan Key Laboratory for Intelligence and Information Technology(哈尔滨工业大学(深圳); 大湾区大学; 东莞市智能信息技术重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对微手势识别中运动易被掩盖及MLLMs处理运动学困难的问题,提出GMoT模块,通过空间加权池等提取运动线索并融合,引入渐进奖励引导策略优化范式,在多数据集上表现出色,还提出BRG召回及跨域转移协议。

AI 中文摘要

微手势识别需要检测短暂的、空间局部化的运动,这些运动经常被占主导地位的静态外观和背景噪声所掩盖。虽然多模态大语言模型(MLLMs)在一般视频理解方面表现出色,但它们在处理微妙的运动学方面存在固有困难,并且通常依赖于静态姿势先验。为此,我们提出了GMoT,一种门控运动感知令牌化模块,它在时间建模之前将稀疏的运动学证据明确提炼成一个紧凑的序列。GMoT通过空间加权池动态突出与动作相关的区域,提取相邻帧的时间差分以捕获精确的运动能量,并使用保守初始化的语义门将这些线索自适应地融合到视觉流中。为了从简单分类过渡到基于证据的推理,我们进一步引入了一种渐进奖励引导的策略优化范式,由一个生成解剖学重点字幕的半监督注释管道支持。除了在iMiGUE(67.32%)和SMG(73.11%)上的比较方法中实现最佳的Top-1准确率,将Qwen3-VL-8B基线提高了+6.80和+3.11个百分点外,我们的框架还引入了基于正确预测的身体区域定位(BRG)召回作为解剖学定位代理,以及iMiGUE和SMG之间的重叠标签跨域转移协议。广泛的评估表明,我们的GMoT增强模型提高了域内准确率,在标签保留损坏下保持了明显的收益,并在明确的小分割条件下提高了面向准确率的跨域转移,同时在其生成的理由中保持了高解剖学定位。

英文摘要

Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. While Multimodal Large Language Models (MLLMs) excel at general video understanding, they inherently struggle with subtle kinematics and often rely on static posture priors. To this end, we propose GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling. GMoT dynamically spotlights action-relevant regions via spatially weighted pooling, extracts adjacent-frame temporal differencing to capture precise motion energy, and adaptively fuses these cues into the visual stream using a conservatively initialized semantic gate. To transition from simple classification to evidence-grounded reasoning, we further introduce a progressive reward-guided policy refinement paradigm, supported by a semi-supervised annotation pipeline that generates anatomically focused captions. Beyond achieving the best Top-1 accuracy among the compared methods on iMiGUE (67.32\%) and SMG (73.11\%), improving the Qwen3-VL-8B baseline by +6.80 and +3.11 points, our framework introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions, together with an overlapping-label cross-domain transfer protocol between iMiGUE and SMG. Extensive evaluations demonstrate that our GMoT-augmented model improves in-domain accuracy, retains clear gains under label-preserving corruptions, and improves accuracy-oriented cross-domain transfer under explicit small-split caveats while maintaining high anatomical grounding in its generated rationales.

CommentsAccepted to ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑