发表机构
Pratt School of Engineering, Duke University; Department of Engineering Science, University of Oxford; Khoury College of Computer Science, Northeastern University(杜克大学普拉特工程学院; 牛津大学工程科学系; 东北大学库里计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出隐式提示机制Hidden-Shot,通过提取隐含视觉任务信息并融合全局任务感知纹理提示,增强低级视觉通用模型在新任务上的一次性泛化能力,并引入C/U评估框架系统测试泛化性能。
AI 中文摘要
尽管低级视觉通用模型引起了广泛关注,但它们在已学习任务之外的零样本/少样本场景中的有效性尚未得到验证。开发理想通用模型的主要挑战在于实现对新未见任务的泛化能力,这也可以通过匹配的定量标准来评估。现有方法在提示工程方面取得了一些进展,但尚未系统地探索这一差距在广泛低级视觉任务中的表现。受此问题启发,我们提出Hidden-Shot,一种隐式提示机制,旨在探索视觉通用模型中的低级任务适应。具体而言,该方法提取隐式视觉任务信息,利用全局任务感知纹理提示,并选择性地将隐式信息与任务内处理信息合并,以增强新任务的一次性能力。整体设计以低成本方式执行直接注入,同时最小程度地改变原始通用模型的架构。此外,我们引入一个数据驱动的评估框架,称为C/U评估,涵盖两种基本场景:3C4U(3个常规任务和4个非常规任务)用于重新训练现有模型,以及3C7U(3个常规任务和7个非常规任务)用于从头训练,作为全面评估系统测试低级通用模型的泛化能力。在七个和十个数据集上的实验分别通过3C4U和3C7U框架验证,性能优于最先进的视觉通用模型。我们提出的Hidden-Shot方法在新任务的一次性场景中表现出优越性能,同时在现有任务上保持一致的性能。
英文摘要
Despite the intense engagement surrounding low-level vision generalist models, their effectiveness in zero/few-shot scenarios beyond learned tasks remains unverified. The primary challenge of developing an ideal generalist lies in achieving the ability to generalize from new unseen tasks, which also can be assessed by matched quantitative criteria. Existing methods have made some progress in prompt engineering but have not systematically explored this gap across a wide range of low-level visual tasks. Stimulated by the problem, we propose Hidden-Shot, an implicit prompt mechanism aimed at exploring low-level task adaptation in a vision generalist model. Specifically, the method extracts implicit visual task-based information, utilizes a global task-aware textural prompt, and selectively merges implicit information with in-task processing information to enhance one-shot capabilities in new tasks. The overall design performs direct injection in a cost-effective manner, while minimally altering the architecture of the original generalist model. Additionally, we introduce a data-driven evaluation framework termed C/U assessment to cover two basic scenarios, 3C4U (3 conventional and 4 unconventional tasks) for retraining existing models and 3C7U (3 conventional and 7 unconventional tasks) for training from scratch, as a comprehensive assessment to systematically test the generalization ability of low-level generalist models. Experiments on seven and ten datasets outperform the state-of-the-art vision generalist model, respectively verified by 3C4U and 3C7U framework. Our presented Hidden-Shot approach demonstrates superior performance on one-shot new tasks while maintaining consistent performance on existing tasks.
CommentsAdded experimental results, corrected minor issues