Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
看见即信仰(并预测):基于视觉语言模型的上下文感知多人类行为预测
机构 * Institute of Architecture of Application Systems, University of Stuttgart, Germany(应用系统建筑研究所,斯图加特大学,德国) ; Bosch Research, Germany(博世研究,德国)
专题命中 VLM训练与架构 :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
AI总结 CAMP-VLM通过结合视觉语言模型与上下文特征,提升了多人类行为预测的准确性,其在预测精度上比基线模型高66.9%。
Comments Accepted at IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026