用于从视频理解复杂装配动作的组合上下文微调视觉语言模型
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
浏览论文内容
中文总结 AI 辅助
针对装配动作理解难题,提出组合上下文微调方法及层划分交替训练方法,将动作分解为语义元素并微调视觉语言模型,创建相关数据集,实验证明该方法优于基线且能提供可解释预测,助力人机协作装配。
中文摘要 AI 辅助
装配动作理解是人机协作装配的关键,但由于微妙动作和精细的手-物体交互而具有挑战性。我们采用组合上下文微调(CCFT)使视觉语言模型(VLMs)适应这一领域,该方法将装配动作分解为语义元素(动词、物体、工具),并使用模板问答对微调VLMs以识别每个动作元素,确保输出具有近确定性。为在有限数据下实现高效多任务学习,提出层划分交替训练(LP-AT)方法,通过特定元素的低秩适配器分配不同模型层识别特定动作元素。此外,我们从现有装配视频数据集创建了HA-ViD-VQA和IKEA-ASM-VQA数据集。实验表明,我们的方法优于强大的动作识别基线,并提供可支持多种下游应用的可解释元素级预测。
英文摘要
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Context Fine-Tuning (CCFT), a method that decomposes assembly actions into semantic elements (Verb, Object, Tool) and fine-tunes VLMs to recognize each action element using templated question-answering pairs. This approach ensures near-deterministic outputs. To enable efficient and effective multi-task learning under limited data, a Layer-Partitioned Alternating Training (LP-AT) method is presented, which assigns distinct model layers to recognize specific action elements through element-specific low-rank adapters. LP-AT alternates weight updates across element-specific adapters, reducing cross-task interference while enabling per-adapter hyperparameter optimization. Furthermore, we create HA-ViD-VQA and IKEA-ASM-VQA datasets from existing assembly video datasets. Extensive experiments on these datasets demonstrate that our method consistently outperforms strong action recognition baselines while providing interpretable element-level predictions that can support diverse downstream applications.
发表机构
- New York University Abu Dhabi(纽约大学阿布扎比分校)
- The University of Auckland(奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。