arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Zero-MELO:基于多模态大语言模型的测试时证据校准用于零样本微手势识别

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

Chengyan Wang, Hanliang Xie, Yueyi Yang, Haoyu Chen

arXiv 2608.14854首次发表:更新:

发表机构

University of Oulu; Peking University(奥卢大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态大语言模型在微手势识别中局部证据不足、分数偏差的瓶颈,提出Zero-MELO框架,结合树搜索、测试时校准与多线索融合,在iMiGUE和MA-52数据集上显著优于Qwen2.5-VL基线。

AI 中文摘要

尽管多模态大语言模型(Multimodal Large Language Models, MLLMs)在通用视频理解方面表现出色,但其在细粒度和以运动为中心的任务中的能力仍有限。这种限制在微手势识别(Micro-Gesture Recognition, MGR)中尤为关键,其中微手势(micro-gestures, MGs)是细微、短时长且空间局部化的人类动作,是隐式情感分析的关键判别信号,但在常见的提示实践中容易被忽略。尽管已有许多判别式方法对MGR进行了深入研究,但将MLLMs用于MGR的探索仍不充分,性能明显较差。我们假设,MLLMs对运动敏感的表征能力受限于其固有的单次前向推理,可通过精心设计的测试时指导大幅提升。受此启发,基于我们先前关于视频大语言模型(Video LLMs)时间不敏感性的研究结果,我们在负对数似然(Negative Log-Likelihood, NLL)空间中分析零样本MGR的错误。我们观察到MLLMs存在两个瓶颈:1)局部证据不足;2)由语言和与运动无关的外观驱动的严重分数偏差。因此,我们提出一种新颖的测试时证据校准框架,以同时提升推理细节和预测可靠性。具体而言,我们引入树搜索机制逐步获取局部、细粒度的视觉证据,并结合测试时校准模块以缓解分数偏差。多线索融合模块随后整合来自多个线索的证据,不依赖单一线索进行最终预测。我们的框架在iMiGUE数据集上的平均类别准确率达到26.84%,在MA-52数据集上达到22.10%,显著优于Qwen2.5-VL基线模型,该基线模型在对应数据集上的准确率分别为16.15%和10.20%。代码将在此httpsURL上提供。

英文摘要

While Multimodal Large Language Models (MLLMs) excel in general video understanding, their capability in fine-grained and motion-centric tasks remains limited. This limitation is particularly critical in micro-gesture recognition (MGR), where micro-gestures (MGs) - subtle, short-duration, and spatially localized human movements - serve as key discriminative signals for implicit affective analysis, yet are easily neglected following common prompting practices. Although MGR has been intensively studied by many discriminative approaches, the use of MLLMs for MGR is underexplored, with notably poor performance. We hypothesize that the motion-sensitive representation ability of MLLMs is constrained by their inherent single-pass forward inference, which can be substantially enhanced through carefully designed test-time guidance. Motivated by this, building on our prior findings regarding temporal insensitivity in Video LLMs, we diagnose zero-shot MGR errors in the Negative Log-Likelihood (NLL) space. We observe that MLLMs suffer from two bottlenecks: 1) insufficient localized evidence and 2) severe score biases driven by language and motion-agnostic appearances. Thus, we propose a novel test-time evidence calibration framework that improves both reasoning details and prediction reliability. Specifically, we introduce a tree search mechanism to progressively acquire localized, fine-grained visual evidence, coupled with a test-time calibration module to mitigate score biases. The multi-cue fusion module then integrates evidence from multiple cues without relying on a single cue for final prediction. Our framework achieves mean-class accuracies of 26.84\% on iMiGUE and 22.10\% on MA-52, significantly outperforming the Qwen2.5-VL baseline, which produces 16.15\% and 10.20\%, respectively. The code will be available at https://zero-melo.github.io/Zero-MELO.

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑