arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你只需重新编程一次:重新思考视觉重新编程的长时间训练

You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming

Zizhao Li, Mohammed Yaqoob Ansari, Xinyu Su, Jiayang Ao, Joseph West, Kourosh Khoshelham

arXiv 2609.36661首次发表:更新:

发表机构

The University of Melbourne; Fudan University(墨尔本大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出YORO方法,通过单次前向遍历利用冻结响应空间构建预测器,无需反向传播,显著提升视觉重新编程效率与准确率。

AI 中文摘要

视觉重新编程是一种参数高效的方法,用于适配预训练模型,然而其训练可能仍然计算昂贵:即使骨干网络被冻结,视觉提示通常也需通过完整模型优化数百个周期。在改变预训练模型所见内容之前,我们询问是否已充分利用它已告诉我们的信息。我们发现,建模完整的源响应可以在无需提示优化的情况下产生强大的下游预测。受此观察启发,我们引入了你只需重新编程一次(YORO),该方法通过单次前向遍历从冻结的响应空间构建下游预测器。其贝叶斯判别映射(BDM)从流式类别统计中推导出协方差感知的仿射映射,无需反向传播、优化器更新或重复访问训练集。当进一步的输入适配有帮助时,YORO-FP可选地将视觉提示微调20个周期。BDM还自然扩展到CLIP,通过将属性-提示相似度视为源响应。在三个全数据设置中,YORO相比最强的先前无梯度映射将平均准确率提高了18.4--24.4%。在16-shot CLIP上,它将四个骨干网络的平均准确率从71.4%提升至77.2%。YORO-FP在选定任务上提供进一步增益,而验证通常保留单遍预测器。这些结果表明视觉重新编程的不同默认方式:先读取冻结响应,仅在需要时优化输入。

英文摘要

Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4--24.4\%. On 16-shot CLIP, it raises the four-backbone average from 71.4\% to 77.2\%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑