发表机构
Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种无训练的多模态大语言模型流水线,动态分配子任务至适配的模型,在2026年MAC微动作挑战赛MA-Bench赛道的开放式任务上获显著性能优势,取得第一名。
AI 中文摘要
微动作是人类在几乎无意识的情况下做出的细微、短暂、低幅度的身体动作,比如坐立不安的手或轻微的头部倾斜,这些动作能可靠地泄露情绪和心理状态。理解微动作不仅是赋予标签,模型还必须描述哪些身体部位在移动,并合理地说明为什么某段视频片段属于特定的细粒度类别。我们提出了一种仅用提示的无训练系统,该系统在2026年MAC~2026微动作挑战赛的细粒度理解赛道(MA-Bench)中获得第一名,该赛道不允许使用微调或真值监督。该系统完全基于冻结的多模态大语言模型(MLLM)构建,将8个子任务动态分配给经实证验证最适合该任务的MLLM:用于封闭式识别任务的判别式MLLM,以及用于开放式描述和推理任务的生成式MLLM。该架构在开放式任务上取得了统计学上显著的性能优势,平均得分达到2.68(5分制),而第二名方法的平均得分为1.44。
英文摘要
Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little conscious intent yet that reliably leak emotional and psychological state. Understanding them goes beyond assigning a label: a model must also describe which body parts move and reason, faithfully, about why a clip warrants a particular fine-grained category. We present the training-free, prompt-only system that won first place in the fine-grained understanding track (MA-Bench) of the MAC~2026 Micro-Action Challenge, where both fine-tuning and ground-truth supervision are disallowed. Built entirely upon frozen multimodal large language models (MLLMs), the system dynamically routes each of the eight sub-tasks to the MLLM empirically best suited for that task: a discriminative MLLM for closed-ended recognition tasks and a generative MLLM for open-ended description and reasoning tasks. This architecture achieves a statistically significant performance advantage on open-ended tasks, attaining an average score of 2.68 (on a five-point scale) compared to 1.44 for the second-best approach.
CommentsAccept at ACM Multimedia 2026