VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
VLM-PAR:一种用于行人属性识别的视觉语言模型
机构 * Department of Innovation Engineering(创新工程系) ; University of Salento, Italy(意大利萨伦托大学) ; Institute of Applied Sciences and Intelligent Systems - CNR(应用科学与智能系统研究所 - CNR) ; University of the Basque Country UPV/EHU(巴斯克国家大学UPV/EHU) ; IKERBASQUE, Basque Foundation for Science(伊基塔斯克巴塞克基金会) ; Sorbonne University Abu Dhabi(索邦大学阿布扎比分校)
专题命中 VLM训练与架构 :VLM(title,abstract);vision language model(title);分类 cs.CV、cs.AI
AI总结 VLM-PAR通过整合大规模视觉语言预训练与跨模态细化,提升行人属性识别在类别不平衡和泛化挑战中的性能。