ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video
ViMoNet: 一种多模态视觉-语言框架,用于从运动和视频中理解人类行为
机构 * Department of Computer Science, American International University–Bangladesh (AIUB)(美国国际大学-孟加拉国计算机科学系) ; Faculty of Psychology, Shinawatra University(信武大学心理学系) ; Faculty of Information Science and Technology, Multimedia University(多媒体大学信息科学与技术系) ; Department of Information Technology, Washington University of Science & Technology(华盛顿科学与技术大学信息科技系)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV
AI总结 ViMoNet通过整合运动和视频数据,提出多模态视觉-语言框架,有效提升人类行为理解与医疗保健应用潜力。
Comments This is the preprint version of the manuscript. It is currently being prepared for submission to an academic conference