LLaVAction: evaluating and training multi-modal large language models for action understanding
LLaVAction:评估和训练多模态大语言模型进行动作理解
专题命中 GUI与屏幕智能体 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结 LLaVAction通过引入动作标记和两阶段流程提升多模态大语言模型的动作理解能力,显著提升基准测试性能。
Comments https://github.com/AdaptiveMotorControlLab/LLaVAction
Journal ref International Conference on Learning Representations (ICLR) 2026