Adapting Vision Foundation Models with Cascaded Semantics
利用级联语义适配视觉基础模型
机构 * University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Missouri–Kansas City(密苏里大学堪萨斯城分校) ; Northeastern University(东北大学) ; Tulane University(杜兰大学) ; University of Bristol(布里斯托大学) ; Oak Ridge National Laboratory(橡树岭国家实验室)
AI总结 该研究针对现有视觉提示调优(VPT)未利用先验知识的问题,提出向VPT注入两类语义先验的级联方案,在34个图像分类数据集上仅调优0.74%的ViT参数即实现优异下游适配效果。
Comments Accepted by Transactions on Machine Learning Research (TMLR), 2026. Project page: https://xixiaouab.github.io/Cascaded-Semantics/